Blog / Recipes / FIG. 139
Community Question Sets Worth Stealing
Real judge prompt examples from public Jev builds: content scorers, deploy gates, code checks. A question set library, with each source linked.
The fastest way to write good judge questions is to read other people's. Builders working with Jev, TypeSafe AI's decision model, have published a surprising number of their question sets in repos, threads and write-ups. This page collects the judge prompt examples worth borrowing, grouped by job, so you can start from a shape that already worked.
The rules for writing questions (operational wording, one judgment per question, probing for bias) belong to the judge questions guide. Here we only curate. Any accuracy or cost figure is as reported by the builder, not checked by us.
Content and growth: the densest sets
Social content is where builders asked the most questions per item, and published them most openly.
- 3,282 posts, eight questions each. Ian Nuttall's set covers topic, hook, tone and whether the post teaches something. A good starter set because it mixes one Choice question (topic) with several yes-or-no checks. Reported: $0.1282 for the run.
- 100,000 viral posts, fourteen yes-or-no questions. "Does the hook open a loop?" "Is there a number in the first line?" "Is the proof real or claimed?" Note how concrete each one is; nothing asks "is this good?"
- Post scoring with SuperX. 61 questions per draft. Rob Hallam says the scorer was fitted on 9,481 posts from 207 creators; the directory entry notes we have not checked those accuracy figures, and neither should you assume them.
- 724 competitor ads, broken down. Hook, format, offer, call to action, awareness stage, and landing-page mismatch. That last question is the one to steal.
Editorial and web review
- Every's editorial vibe check. 21 questions per document across 37 documents. Mike Taylor planted seven defects; Jev found six, and only the larger model found the seventh, as reported. That's the honest way to test a question set: plant known answers.
- AI slop detector. 35 tells, from purple gradients to fake testimonials. A good template for any checklist-style audit where each tell is its own yes-or-no.
- Jev Web Analyzer. Ten bounded questions about first-visit comprehension of a SaaS page, with fetching, caching and presentation kept deterministic.
- Doomscroll Filter and the X timeline labeler. Eight questions that sort posts into Read, Skim or Pass; five labels (clean, engagement bait, promo, secondhand, filler). Proof that a small, well-named label set beats a long one.
Engineering gates
The sets builders put in front of risky actions. Note that none of them auto-approve on a single verdict.
- HeyJev deploy approval demo. 18 typed questions about a deployment, including tests, rollback and availability, turned into an inspectable verdict.
- jev-gates. Seven gates for Claude Code: rules, scope, intent, done, claims, proof and commit honesty. They escalate but never approve.
- tripwire. Seven checks on every LLM response, asked in one call.
- jev-lint. Does the function do what its name says? Is the comment still true? Does the test verify what it claims? Three questions that no regular linter can ask.
For the engineering side of wiring questions into agent hooks, see the jev-engineering guide.
Libraries to browse
- TypeSafe AI playground. 110 use cases, games and dilemmas with editable prompts and A/B tests. The closest thing to a question set library the community has.
- Awesome Jev Skills. Nine installable agent skills with a catalogue of scenarios to copy.
- Jev plays Puyo Puyo. A small but instructive set: one question about the game phase, three about which outcome fits the strategy. Board-only questions were too weak, per the author, so he staged them.
How to steal responsibly
Copy the structure, not the confidence. A set that worked on someone's posts is a hypothesis on yours. Run it against a small hand-labeled sample first, as building gold sets describes, and version your questions once they're in use, per question versioning. Every set above was tuned on its author's data; your threshold will differ.
Frequently asked questions
Where can I find real LLM judge prompt examples?
The builds above publish theirs in repos, threads or write-ups; each entry links to the original source. The TypeSafe AI playground's 110 use cases is the broadest single collection we've cataloged.
How many questions should a judge set have?
Published sets here range from four questions per move to 61 per draft. Start small and add questions only when a gold set shows a gap, as the judge questions guide recommends.
Can I reuse someone else's question set as-is?
You can start from it, but validate it on your own labeled data before trusting its verdicts. Thresholds and wording that worked for one dataset rarely transfer untouched.
Do these question sets work with models other than Jev?
The questions are plain language, so the shape transfers to any judge model. Costs and calibration will differ, so re-measure both.
Numbers throughout are as reported by the build authors, not verified by shipwithjev. Code-shaped examples are pseudocode; the official docs live at docs.typesafe.ai.