Blog / Recipes / FIG. 81
Build Your Team's Pre-Publish Content Scorer
Build an AI content scoring check with Jev: turn your style guide into yes/no questions, score every draft in seconds, keep editors in charge.
Every content team has a style guide, and every content team has drafts that ignore it. AI content scoring turns that guide into a check that runs on every draft before an editor sees it: a battery of closed questions, each with a verdict and a confidence. Jev, TypeSafe AI's decision model, is fast and cheap enough to run it on every save, not just before publish.
How pre-publish scoring fits alongside teardowns, reply triage, and the rest of the marketing stack is covered in Jev for marketers. This page builds the scorer.
Two receipts worth copying
Rob Hallam's SuperX scorer asks 61 questions about a draft post in about a second, for $0.0004, as reported (build). He also says it was fitted on 9,481 posts and picks the more viral of two posts two times in three; those accuracy figures are his and haven't been independently checked.
Mike Taylor ran Jev over 37 documents he'd written for Every, 21 questions each: 1,709 judgments for under a cent, with a median of 0.35 seconds per passage against 8.83 seconds for Fable 5.1, as reported (build). Jev found six of the seven defects he'd planted; only the larger model found the seventh. That's the honest shape of it: fast and cheap on the first pass, not infallible. No official Jev benchmarks exist.
The recipe
-
Extract rules from your style guide (45 minutes). Go through it and pull out every rule that can be judged from the draft alone. "Be engaging" isn't a rule. "The first paragraph states what the reader will get" is.
-
Turn each rule into a yes/no question. Examples:
- "Does the first paragraph state a concrete benefit or outcome for the reader? YES / NO / UNCLEAR"
- "Does the draft make a numerical claim without naming a source? YES / NO / UNCLEAR"
- "Does the headline promise something the body does not deliver? YES / NO / UNCLEAR"
- "Does the draft use jargon without defining it for a non-specialist? YES / NO / UNCLEAR"
Writing these well is the whole game. How to write judge questions is required reading here, especially on avoiding compound and vague questions.
-
Weight and group them. Mark each question as blocker (fact-checking, legal, brand safety), major, or minor. The score an editor sees is "2 blockers, 1 major", not a single number from 0 to 100.
-
Wire it into the draft flow.
# pseudocode, not real API syntax
on draft_saved(d):
results = [(q, jev.choose(q.text, d.body, ["YES", "NO", "UNCLEAR"])) for q in STYLE_QUESTIONS]
flags = [r for r in results if r.fails(q.pass_answer) or r.low_confidence()]
show_panel(d, group_by_severity(flags))
Per ecosystem documentation, the model is typesafe-ai/jev via the Vercel AI Gateway with per-choice probabilities. Syntax at docs.typesafe.ai. Drafts longer than a few thousand words may need to be checked in sections; check current input limits in the docs.
-
Calibrate on old posts (1 hour). Run the scorer over 20 published pieces your editors already consider good and 20 they consider weak. If it flags the good ones heavily, questions are too strict or too vague. Fix wording before anyone sees a live panel.
-
Launch as advisory. Writers see flags; editors still approve. After a month, look at which flags editors consistently agreed with. Those can become required checks.
A draft quality checker, not a ghostwriter
The obvious objection: "Why not have AI just fix the draft?" Because Jev doesn't write. It tells you the intro lacks a benefit; it can't write a better intro. If you want suggested rewrites, trigger a chat model only on flagged sections, and let the writer decide. Verdicts from the decision model, prose from the generative model or the human.
The other objection: "Won't writers game it?" Some will, the same way they game any checklist. That's fine if the checklist encodes things you actually want. Writing to pass "names a source for every number" is a good outcome.
Keep editors in charge
The scorer never blocks publishing on its own. Publishing is public and hard to take back. A blocker flag means an editor must look, not that the system refuses.
Version the question set. When the style guide changes, update the questions in one place and note the date, so month-over-month scores stay comparable.
Don't use scores to rank writers. A quality checker used as a performance metric turns into a gaming target fast, and it was never designed to evaluate people.
Frequently asked questions
How fast is AI content scoring with Jev?
Builders report around a second for 61 questions on a post, and a 0.35-second median per passage in another run, both as reported. Long drafts checked in sections will take proportionally longer.
Can Jev rewrite my draft?
No, it only returns verdicts. Pair it with a chat model for suggested rewrites on flagged sections, and keep the writer in control of what changes.
How many questions should a content scorer ask?
Start with 10 to 20 drawn from your style guide, then add more as editors spot recurring issues. Builders have run 61 per draft, as reported, but more questions only help if each is clearly worded.
Is a pre-publish scorer useful for social posts too?
Yes, and that's where the SuperX build started. See Jev for marketers for how teams apply scoring across channels.
Numbers throughout are as reported by the build authors, not verified by shipwithjev. Code-shaped examples are pseudocode; the official docs live at docs.typesafe.ai.