shipwithjev

Blog / Builds & people / FIG. 63

Jev's Context Window: What Fits in a Question (And What Shouldn't)

How much text fits in a Jev question, why the honest answer is "less than you want, more than you need," and the section-and-aggregate pattern.

The number you came for lives at docs.typesafe.ai, and we don't republish expiring specs. The design truth that outlasts any number: decision models reward small inputs by construction, and pipelines that fight that grain lose accuracy before they ever hit a hard limit.

Why small wins even when big fits: a verdict's quality tracks how directly the evidence bears on the question, and padding context dilutes exactly that. The craft guide's rule, one judgment per question over the minimum evidence that decides it, is also the performance advice; the 61-questions-per-draft build sends a post, not a corpus, per question, and that's why it's fast, cheap, and sharp at once. Cost follows the same slope: per-verdict pricing moves with context size, so lean questions stay at the cheap end of the band.

For genuinely large inputs, contracts, transcripts, long threads, the pattern is settled and lives in document classification: section, judge, aggregate. Split by natural structure, run per-section verdicts, and roll up with explicit logic ("flag the document if any section flags"), which also buys you the debuggability a monolithic judgment never offers: when something fires, you know which page fired it. Whole-document reasoning, synthesis across everything at once, was never this model's job anyway; that's the frontier tier of your cascade, used sparingly, where its context appetite earns its price.

The anti-pattern to name: stuffing context to avoid designing questions. If you're tempted to send everything "so it has what it needs," the question isn't specified yet, and that's the actual work.

Frequently asked questions

What's the actual token limit?

The official docs' current figure; it can change and we won't chase it. Design for small and the limit stops being load-bearing.

My input is a 40-page document. Now what?

Section-judge-aggregate per the document pattern: per-section verdicts with explicit roll-up logic, frontier reasoning reserved for the rare whole-document questions.

Does more context improve verdicts?

Relevant context does; volume doesn't, and usually hurts. The evidence that decides the question, no more, is both the accuracy and the cost optimum.

Numbers throughout are as reported by the build authors, not verified by shipwithjev. Code-shaped examples are pseudocode; the official docs live at docs.typesafe.ai.