0449GitHub
jev-korean-benchmark
Small Korean/English sample study with recorded responses, including medical-text questions; not a clinical validation.
mahlernim/jev-korean-benchmarkREADME ↗
# Jev in Korean: a 100-question sample check **[한국어로 읽기 ↓](#jev의-한국어-성능-100문항-표본-점검)** Can you use TypeSafe Jev on Korean text, or should you translate everything to English first? This is a small, frozen, reproducible check that tries to answer that — 100 questions per cell, drawn from four public test sets, with every response recorded. It is a sample check, not a benchmark: 100 questions give roughly ±8 points of uncertainty, so small differences here are not findings. ## Short answer - **Korean reading comprehension is fine.** 96 of 100 in Korean against 97 in English, on the same questions. There is no meaningful cost. - **Fine-grained meaning judgements are shaky in both languages.** 76 in Korean, 80 in English. The Korean number is a little lower, but 80% is not good in *either* language — this is the weak spot regardless of which one you use. - **Don't bother writing your prompts in English.** Korean content scored within a point or two whether the instructions were Korean or English. Translating your instructions buys nothing. - **The two medical numbers cannot be compared to each other.** 89 on an English exam and 80 on a Korean one are two different tests, not evidence about language. See [About the medical numbers](#about-the-medical-numbers). - **Reordering your input changes about 1 answer in 8**, equally in both languages. Worth knowing, but it is not a Korean-specific problem. Bar ends are the score; the connector shows the gap between English and Korean on identical questions. Medical rows sit below the line as single points because no counterpart exists for them. ## What was tested <!-- results-table-start --> | Task | Language | Jev | Luna | Luna − Jev (95% CI) | |---|---|---:|---:|---:| | Belebele reading | English | 97/100 | 98/100 | +1 (−3