shipwithjev

Blog / Questions / FIG. 56

Does Jev Work in Languages Other Than English? What Reports Show

Can Jev judge non-English text? What ecosystem reports and builds suggest about multilingual verdicts, plus the calibration rule that matters per language.

The honest tier-list of what's known: TypeSafe hasn't published a formal language-support matrix we can cite, ecosystem usage is overwhelmingly English so the evidence base is thin, and underlying language-model architectures generally carry substantial multilingual capability by default. Translation: probably yes for major languages, verify for yours, and nobody gets to be confident yet, including us.

What tips the "probably": the model family Jev belongs to inherits multilingual pretraining as a rule, and closed-set judgment is actually the easier multilingual task, picking among fixed choices demands less generative fluency than composing text would, which is why classification has always traveled across languages better than generation. The structural advice that follows: keep your questions and choice labels in English (the instruction-following backbone) while the evidence text stays in its source language, which is the pattern multilingual judge pipelines converge on generally and costs nothing to adopt.

What tips the "verify": calibration is per-language, full stop. The hour-long ritual, 100 labeled cases, agreement scored, runs per language you serve, because accuracy in English predicts accuracy in Spanish about as well as vibes do, and confidence thresholds may want different settings where the model is less sure. Slang, dialect, and code-switching (the real texture of moderation and support streams) deserve deliberate cases in that set. And the fallback is already designed: low-confidence non-English verdicts route up the cascade to a frontier model whose multilingual depth is documented, exactly like any other hard slice.

If you run multilingual verdicts at volume, submit the numbers: this is the ecosystem's thinnest evidence base, and your calibration table would be the most useful receipt on the site.

Frequently asked questions

Which languages does Jev officially support?

No official matrix exists at the time of writing; docs.typesafe.ai is where one would appear, and this page updates when it does.

Should I translate text before judging it?

Usually no: translation adds cost, latency, and its own errors before the verdict even runs. English questions over source-language evidence, calibrated per language, is the cleaner default.

What about right-to-left and CJK languages?

Same answer with sharper emphasis: plausible by architecture, unproven by public receipts, calibrate before trusting, and escalate low confidence rather than forcing it.

Numbers throughout are as reported by the build authors, not verified by shipwithjev. Code-shaped examples are pseudocode; the official docs live at docs.typesafe.ai.