Runs & evidence
Loading runs…
Index
Skipped or incomplete runs
Limitations
- KoBALT-700 is a public benchmark, so prior exposure of these items in model training (contamination) is unknown and cannot be ruled out.
- One run per configuration: there is no measure of run-to-run repeatability, and close scores may not reflect a real difference.
- The recorded dataset revision is
0.0.0, a placeholder — it is not content-addressed, so exact dataset identity rests on the benchmark release rather than a pinned hash. - Reasoning policies differ across runs (provider default, explicit minimal/low effort, disabled); comparisons across those settings are about configuration, not models alone.
- Small subgroups are noisy: several subclasses have fewer than 20 items, and single items can move those percentages by several points. Treat domain and level breakdowns as descriptive.