KoBALT-700 benchmark record

Independent evaluation runs on a Korean linguistics benchmark, with evidence.

Runs & evidence

Loading run metadata…

Loading runs…

Index

Skipped or incomplete runs

Limitations

  • KoBALT-700 is a public benchmark, so prior exposure of these items in model training (contamination) is unknown and cannot be ruled out.
  • One run per configuration: there is no measure of run-to-run repeatability, and close scores may not reflect a real difference.
  • The recorded dataset revision is 0.0.0, a placeholder — it is not content-addressed, so exact dataset identity rests on the benchmark release rather than a pinned hash.
  • Reasoning policies differ across runs (provider default, explicit minimal/low effort, disabled); comparisons across those settings are about configuration, not models alone.
  • Small subgroups are noisy: several subclasses have fewer than 20 items, and single items can move those percentages by several points. Treat domain and level breakdowns as descriptive.