KoBALT-700 benchmark record

Independent evaluation runs on a Korean linguistics benchmark, with evidence.

Results

Loading run metadata…

Loading results…

Comparison, by reasoning group

Runs are grouped by provider reasoning setting because the setting changes scores substantially. The groups are not directly comparable with each other; only rows within the first table share identical conditions. Accuracy is strict exact match (chance: 10%).

Provider-default reasoning

Reasoning-limited (explicit minimal/low effort)

Reasoning disabled

By domain

Five linguistic domains. Small domains are noisier (Morphology: 42 items; Syntax: 300). Charts show provider-default runs; tables include reasoning-limited runs, marked.

Provider-default runs. Group labels include item counts.
Exact domain values

By difficulty

Level 1: 182 items; Level 2: 220; Level 3: 298. Charts show provider-default runs; tables include reasoning-limited runs, marked.

Provider-default runs. Group labels include item counts.
Exact difficulty values

Reasoning comparisons

Paired runs of the same model under different reasoning settings. Unpaired observations are labeled as such — they show where a configuration landed, not what the setting changed.

Reliability

Invalid means no single A–J choice could be extracted; invalid outputs count as wrong. Latency depends on the endpoint and deployment behind each run and should not be read as intrinsic model speed.