KoBALT-700 benchmark record

Independent evaluation runs on a Korean linguistics benchmark, with evidence.

Overview

KoBALT-700 is a set of 700 expert-written Korean multiple-choice items across five linguistic domains (Syntax, Semantics, Pragmatics, Phonetics/Phonology, Morphology). This record reports our own evaluation runs against it: what was tested, how, and what came out.

Loading run metadata…

What the runs show

  • Findings load with the data.

Overall accuracy, provider-default runs

Strict exact match with 95% Wilson intervals; ten options per item, so chance is 10%. Only provider-default reasoning runs are plotted here — the comparable set. The full comparison, including reasoning-limited runs, is on the Results page.

Provider-default runs only. Axis starts at zero.

Further reading

Limitations, in brief

  • KoBALT-700 is public, so prior exposure of these items in training cannot be ruled out.
  • One run per configuration; close scores may not reflect a real difference.
  • Reasoning settings differ across runs — cross-setting comparisons are about configuration, not models alone. Full list on the Runs & evidence page.