Overview
KoBALT-700 is a set of 700 expert-written Korean multiple-choice items across five linguistic domains (Syntax, Semantics, Pragmatics, Phonetics/Phonology, Morphology). This record reports our own evaluation runs against it: what was tested, how, and what came out.
What the runs show
- Findings load with the data.
Overall accuracy, provider-default runs
Strict exact match with 95% Wilson intervals; ten options per item, so chance is 10%. Only provider-default reasoning runs are plotted here — the comparable set. The full comparison, including reasoning-limited runs, is on the Results page.
Further reading
- Results — tiered comparison tables, domain and difficulty breakdowns, reasoning comparisons, reliability.
- Method — decoding settings, the exact Korean prompt, extraction and scoring, reproduction.
- Runs & evidence — per-run archive with config and results files, skipped runs, full limitations.
- Upstream: paper (arXiv:2505.16125) · benchmark repository · dataset.
Limitations, in brief
- KoBALT-700 is public, so prior exposure of these items in training cannot be ruled out.
- One run per configuration; close scores may not reflect a real difference.
- Reasoning settings differ across runs — cross-setting comparisons are about configuration, not models alone. Full list on the Runs & evidence page.