Model Bake-off — Real-Task Scorecard

Six models · two phases · four live backlog tasks · judged blind by Codex against audited answer keys · Tue 21 Jul 2026
6models
4real tasks
26agent runs
3fixes live
$13OpenRouter spend

Phase 1 — Research (out of 200)

Each model independently investigated all four pending tasks against identical code packs: is it still needed, root cause, approach, risks. Scored per task on accuracy, verdict, depth, approach, and epistemic honesty against a live-audited answer key. Tap a row for its per-task scores.

Phase 2 — Build (out of 160)

Each model implemented all four fixes in isolated sandboxes from the same verified problem statement. Diffs judged blind against pristine code: correctness, safety, minimality, handoff. Ship flags are the judge's per-patch verdicts; small gaps between neighbouring totals are judge-run noise. Tap a row for per-task scores and flags.

SHIP judge would deploy after reviewDO NOT SHIP disqualifying defect found

Wall-clock per task

Same harness per lane, so within a task the spread is the model, not the pipe. Frontier models don't stream faster — they take fewer, better turns. K3 reasons silently for minutes at a stretch and pays for it in wall time, not correctness.

What actually shipped

TaskWinning patchStatus
Gmail CLI safety (status banner, thread-ID guard, draft-delete)Codex, Fable-reviewedLIVE commit 9277909, live-smoked
Session-monitor OpenRouter spend pillCodex, Fable-reviewedLIVE v0.5.59 + pm2 collector
Cascade preset=180d API fixFableSTAGED branch bakeoff-fixes
KMZ API credits security (idempotency + item ownership)Codex mechanism + 3 correction roundsSTAGED DEPLOY-SAFE, awaiting gate

The two staged fixes share one branch in the kmz-api repo — one deploy decision covers both. Verdict after a four-round adversarial loop: seven findings round 1, zero by round 4.

Process findings — the real yield

01 Dual benchmarks caught a wrong "truth"

Fable and Codex researched independently as the answer-key pair. They disagreed once — Codex said "close the Cascade task", Fable found the bug alive in the API half. The code sided with Fable. A single benchmark would have shipped the error as ground truth.

02 Blind self-judging held up

Codex judged pools containing its own anonymised work and repeatedly ranked itself down — flagging its own Cascade patch DO-NOT-SHIP and placing its own research 5th. No detectable self-preference.

03 The answer key itself carried an error

The original task note framed members as client-owned. Every model inherited it; the audit verified mechanics, not domain semantics. Only the domain owner caught it. Eval oracles are not a substitute for owner review.

04 The winning patch was not deployable

The best build still needed four adversarial rounds (7 findings → 2 → 2 → 0) plus a domain-model correction before earning DEPLOY-SAFE. "Won the bake-off" and "safe for production" are different bars.

05 Cost honesty

Research + builds cost ~$13 of OpenRouter credit — five times the per-round budget. Credit now sits at $22.6 of $25; top up before the next open-model round. Frontier legs ran on subscriptions.