Results
Loading results…
Comparison, by reasoning group
Runs are grouped by provider reasoning setting because the setting changes scores substantially. The groups are not directly comparable with each other; only rows within the first table share identical conditions. Accuracy is strict exact match (chance: 10%).
Provider-default reasoning
Reasoning-limited (explicit minimal/low effort)
Reasoning disabled
By domain
Five linguistic domains. Small domains are noisier (Morphology: 42 items; Syntax: 300). Charts show provider-default runs; tables include reasoning-limited runs, marked.
Exact domain values
By difficulty
Level 1: 182 items; Level 2: 220; Level 3: 298. Charts show provider-default runs; tables include reasoning-limited runs, marked.
Exact difficulty values
Reasoning comparisons
Paired runs of the same model under different reasoning settings. Unpaired observations are labeled as such — they show where a configuration landed, not what the setting changed.
Reliability
Invalid means no single A–J choice could be extracted; invalid outputs count as wrong. Latency depends on the endpoint and deployment behind each run and should not be read as intrinsic model speed.