Explicit Edit / Public benchmark
What changes when the harness changes?
A public leaderboard comparing how models and harnesses behave on precise text-editing tasks. Explore accepted runs, inspect the evidence, and see where the same model performs differently.
Built from community runs. to help us compare more models and harnesses.
Latest additions
From the published dataset- Loading runs…
Leaderboards
Full runs only. Each model × reasoning × harness uses its newest complete version per provider.
Each model × reasoning has its own rating. Combination score is the median of its runs after exact-route version selection. Model score is the median of harness-combination scores. Harness score is baseline-adjusted against baseline-agent routes; missing routes contribute zero change.
Model route × harness
How to read this matrix
Rows are model × provider × reasoning routes; columns are harnesses. Low, high, and other recorded reasoning modes are separate configurations. Missing reasoning is unspecified, not off. Base-model groups start collapsed and reveal each provider × reasoning variant. A collapsed group cell is a display median across its exact routes, not a separate model rating. Single variants stay direct. A compact harness column uses the newest complete version separately for each exact route; repeats contribute a median score. Use a harness arrow to inspect multiple versions. In Score, Across all models shows the baseline-adjusted mean over the fixed set of full baseline-agent model–provider–reasoning routes. Only exact reasoning matches contribute measured differences. Missing routes contribute zero change in totals, never invented source values. Change and measured coverage appear under the score. Filters and the baseline-only switch cannot change this rating. Other metrics show medians across measured model families. Only full runs contribute to summaries. In Score, an incomplete cell shows a dash and the largest single run's task count (N/total tasks) instead of a misleading score. Select a value to compare configurations, or select a row or column label to open details.
Price × score frontier
Solid green: Pareto frontier · Dashed blue: same combination at different reasoning levels · Whisker: repeated runs
Only full runs with a fully measured value on the chosen axis appear; prices are never estimated. Each square is one model route, harness version and reasoning level. Repeated full runs contribute one median point and a whisker if scores differ. The frontier is recalculated for the selected axis. Click a square or frontier entry to compare up to four; use Nearby points when squares overlap. Filters apply here too.
Task families
Tool usage
Runs
Sort or filter accepted configurations. Scroll within the table for more runs; select a cell to see its sources.
Columns
Drag or use the arrows to reorder. Clear a checkbox to hide a column.
How to read this table
Trials are benchmark task executions. Coverage is the number of unique tasks observed out of the benchmark total. Score weights first-attempt exactness more heavily than success after recovery. Select any cell to see how its value was computed, or use the filter button in a column header.