Blog / Comparisons / FIG. 112
Batch vs Real-Time Verdicts
Batch LLM processing or real-time verdicts? Pick by when the answer must exist, how cost scales with clock or items, and what builders report.
The same question ("is this post worth reading?", "which queue does this ticket belong in?") can be asked in two very different ways: in bulk, over a pile of items, or live, the instant each item appears. Batch LLM processing and real-time verdicts use the same model and the same question, and they differ in almost everything else: cost shape, failure modes, and what "done" means. This page owns the split. Examples use Jev, TypeSafe AI's decision model, because its reported speed makes both modes practical, which is exactly why teams now have to choose on purpose.
Scope: quotas, 429s, and throttling are Jev rate limits. What real-time loops look like at the extreme is Jev in real-time games.
Batch LLM processing or real-time: the one question that decides it
When does the verdict need to exist? If something downstream acts on the answer within the same interaction (a move, a route, a block, a UI update), it's real-time. If the answer feeds a report, a dataset, a dashboard, or tomorrow's queue, it's batch. Most confusion comes from teams running batch-shaped work through a live path because the live path already existed.
Two receipts, two cost shapes
Real-time cost scales with the clock. TypeSafe's launch demo played Doom at about ten queries a second for about $7 an hour, as reported by the engineer who built it. Napkin math from those reported figures: roughly 36,000 decisions an hour, around two hundredths of a cent each. The per-decision price is tiny, but a real-time loop pays it every tick, whether or not anything interesting happened. Leave it running for a month and the clock, not the workload, writes the bill.
Batch cost scales with the items. Ian Nuttall ran 3,282 of his posts through eight questions each: about 26,000 verdicts, $0.1282, 8 minutes 34 seconds for the full run, as reported. That job has a fixed size and a fixed price, and when it's done, it's done.
Don't compare the per-verdict figures across those two receipts; the inputs and questions are very different. Compare the shapes. Real-time is a rate times a duration. Batch is a count times a price.
When real-time earns its keep
The answer changes the next moment. Games, agent loops, and live interfaces, where a late verdict is a verdict that didn't arrive. The directory's real-time builds run on sub-second answers: predictive spreadsheets rate each row in about 100 ms as you type the column header, as reported.
The input only exists now. Streams (chat, comments, sensor feeds) where storing and processing later means the moment is gone.
The action is a gate. Screening input before a model or user sees it, or checking an agent's action before it executes, has to happen inline, by definition.
Design costs of real-time: a latency budget per call, a degrade path when the API is slow or unreachable, and deduplication so a burst of identical events doesn't become a burst of identical verdicts.
When batch wins
The answer is read later. Analytics, labeling, backfills, audits, and eval suites. Nobody needs last quarter's tickets classified in 200 milliseconds.
You want to query verdicts, not recompute them. Judging each row once and storing the result as a column turns repeated questions into free lookups. A DuckDB extension that classifies table rows reported about ten seconds per thousand rows; run that once, store it, and every dashboard reads the column. The recipe for this is materializing verdict columns.
You want checkpoints and retries. Batch jobs can resume from where they failed, run at whatever pace your limits allow, and be re-run under a new question version as a clean experiment. Whether TypeSafe offers any batch-specific endpoint or pricing is for docs.typesafe.ai to say; this page doesn't assume one exists.
Streaming vs batch AI: the hybrid most teams end up with
The common production shape is both. Real-time verdicts on new items as they arrive, for the actions that can't wait. A batch job that backfills history, re-judges items when the question version changes, and fills in anything the live path dropped during an outage. And materialized verdicts so reads never trigger new calls. The live path stays small and fast; the batch path does the bulk of the work at its own pace.
Two rules keep the hybrid sane. Tag every verdict with its question version so live and backfilled results are comparable. And never let a batch job share the live path's quota unmanaged; that's how a backfill takes down triage, which is the rate-limits page's territory.
Frequently asked questions
Should I use batch or real-time LLM processing?
Ask when the verdict must exist: if something acts on it within the same interaction, go real-time; if it feeds reports, datasets, or later queues, batch it. Most products end up with both.
Is real-time more expensive than batch?
Per verdict, not necessarily; the difference is how cost scales. Real-time loops pay per tick for as long as they run, as the Doom receipt's roughly $7 an hour shows, while batch jobs pay once per item.
Can Jev run fast enough for real-time use?
Builders report about ten decisions a second in Doom and around 100 ms per spreadsheet rating, but there are no official latency benchmarks. The real-time games page collects the receipts.
Does Jev have a batch API?
Check docs.typesafe.ai for anything official. Builders run batch work today as ordinary jobs with checkpoints, off the live path.
Numbers throughout are as reported by the build authors, not verified by shipwithjev. Code-shaped examples are pseudocode; the official docs live at docs.typesafe.ai.