shipwithjev

Catalog / Research & data

0449GitHub

jev-korean-benchmark

Small Korean/English sample study with recorded responses, including medical-text questions; not a clinical validation.

mahlernim/jev-korean-benchmarkREADME ↗
# Jev in Korean: a 100-question sample check

**[한국어로 읽기 ↓](#jev의-한국어-성능-100문항-표본-점검)**

Can you use TypeSafe Jev on Korean text, or should you translate everything to English first? This is a small, frozen, reproducible check that tries to answer that — 100 questions per cell, drawn from four public test sets, with every response recorded. It is a sample check, not a benchmark: 100 questions give roughly ±8 points of uncertainty, so small differences here are not findings.

## Short answer

- **Korean reading comprehension is fine.** 96 of 100 in Korean against 97 in English, on the same questions. There is no meaningful cost.
- **Fine-grained meaning judgements are shaky in both languages.** 76 in Korean, 80 in English. The Korean number is a little lower, but 80% is not good in *either* language — this is the weak spot regardless of which one you use.
- **Don't bother writing your prompts in English.** Korean content scored within a point or two whether the instructions were Korean or English. Translating your instructions buys nothing.
- **The two medical numbers cannot be compared to each other.** 89 on an English exam and 80 on a Korean one are two different tests, not evidence about language. See [About the medical numbers](#about-the-medical-numbers).
- **Reordering your input changes about 1 answer in 8**, equally in both languages. Worth knowing, but it is not a Korean-specific problem.



Bar ends are the score; the connector shows the gap between English and Korean on identical questions. Medical rows sit below the line as single points because no counterpart exists for them.

## What was tested

<!-- results-table-start -->
| Task | Language | Jev | Luna | Luna − Jev (95% CI) |
|---|---|---:|---:|---:|
| Belebele reading | English | 97/100 | 98/100 | +1 (−3

Also filed under Research & data

  1. 0607

    Verify: new-hire onboarding completion

    An agent reports onboarding done; the judge verifies access and equipment claims.

    everyai-com · Research & data

  2. 0606

    Triage: vague meeting request gets a disposition

    A vendor asks for 30 minutes with no agenda; the judge picks the disposition.

    everyai-com · Research & data

  3. 0605

    Triage: data-loss bug gets a severity

    A note-taking app silently drops edits on flaky networks; the judge grades severity.

    everyai-com · Research & data

  4. 0604

    Triage: crash report routing + reproducibility

    A crash report with steps and logs; the judge routes it and checks reproducibility.

    everyai-com · Research & data