shipwithjev

Blog / Questions / FIG. 121

Is Jev Deterministic?

Is Jev deterministic? No: the same input can draw a different verdict. Why, how much it matters, and the cache-and-log rule that fixes replay.

No. Jev, the decision model from TypeSafe AI, can occasionally return a different verdict, or different per-choice probabilities, for the same input asked the same question twice. If your system needs identical answers on replay, you build that property around the model; you don't get it from the model.

That sounds worse than it is. Here's how much it matters and what builders do about it.

How much the same input can move

Most runs agree, and the variation concentrates where you'd expect: on inputs that are genuinely ambiguous, where the probabilities were already split. A verdict at 0.97 rarely flips. A verdict at 0.52 is a coin the model is telling you it can't call, and asking again is just flipping it again.

For context, one comparison suggests Jev is steadier than chat-model judges: LangChain reported its quality-score variance at 92 to 913 times lower than GPT-5.6 Luna, Terra, and Claude Sonnet 4.6 on continuous scoring, as reported. Lower variance is not zero variance, and there are no official Jev benchmarks on this. The limitations page lists nondeterminism among the soft edges, alongside the rest of the list we won't repeat here.

The cache-and-log rule

The fix is architectural and simple: anything that must replay reads stored verdicts, never re-asks the model.

  • Log every verdict with its input, question version, full probabilities, and timestamp. Audits, datasets, and anything billing-adjacent read from the log.
  • Cache repeat inputs, with the question version in the cache key so a rewording never serves stale rulings. The jevcache library is built for exactly this, as reported: a local decision cache keyed on model, schema, and state, for cheaper repeats and deterministic replay in CI.
  • Gate on confidence so low-probability verdicts escalate instead of flip-flopping silently. The probability outputs explainer covers reading the scores.
  • Never re-roll for a better answer. Asking again until the verdict you wanted appears converts uncertainty into fake confidence.

The production guide covers the wider ops picture, including retries that separate transport failures from low-confidence verdicts. For how a decision model differs from a chat model in the first place, start there.

Frequently asked questions

Why does the same input give a different output?

Hosted models don't guarantee identical outputs across calls, and ambiguous inputs sit near decision boundaries where small variation changes the winner. Clear inputs with sharp questions rarely flip.

Can I make Jev deterministic?

Not the model itself, but your system can be: log verdicts, serve repeats from a cache keyed on question version, and replay from logs. Libraries like jevcache package that pattern.

Does nondeterminism matter for my use case?

For routing a ticket, barely. For audits, reproducible datasets, and compliance, a lot, which is why those read from stored verdicts rather than fresh calls.

Should I set a temperature to fix it?

Check docs.typesafe.ai for what the API actually exposes; we don't document parameters here. Either way, cache-and-log is the reliable fix.

Numbers throughout are as reported by the build authors, not verified by shipwithjev. Code-shaped examples are pseudocode; the official docs live at docs.typesafe.ai.