Each model independently investigated all four pending tasks against identical code packs: is it still needed, root cause, approach, risks. Scored per task on accuracy, verdict, depth, approach, and epistemic honesty against a live-audited answer key. Tap a row for its per-task scores.
Each model implemented all four fixes in isolated sandboxes from the same verified problem statement. Diffs judged blind against pristine code: correctness, safety, minimality, handoff. Ship flags are the judge's per-patch verdicts; small gaps between neighbouring totals are judge-run noise. Tap a row for per-task scores and flags.
Same harness per lane, so within a task the spread is the model, not the pipe. Frontier models don't stream faster — they take fewer, better turns. K3 reasons silently for minutes at a stretch and pays for it in wall time, not correctness.
| Task | Winning patch | Status |
|---|---|---|
| Gmail CLI safety (status banner, thread-ID guard, draft-delete) | Codex, Fable-reviewed | LIVE commit 9277909, live-smoked |
| Session-monitor OpenRouter spend pill | Codex, Fable-reviewed | LIVE v0.5.59 + pm2 collector |
| Cascade preset=180d API fix | Fable | STAGED branch bakeoff-fixes |
| KMZ API credits security (idempotency + item ownership) | Codex mechanism + 3 correction rounds | STAGED DEPLOY-SAFE, awaiting gate |
The two staged fixes share one branch in the kmz-api repo — one deploy decision covers both. Verdict after a four-round adversarial loop: seven findings round 1, zero by round 4.
Fable and Codex researched independently as the answer-key pair. They disagreed once — Codex said "close the Cascade task", Fable found the bug alive in the API half. The code sided with Fable. A single benchmark would have shipped the error as ground truth.
Codex judged pools containing its own anonymised work and repeatedly ranked itself down — flagging its own Cascade patch DO-NOT-SHIP and placing its own research 5th. No detectable self-preference.
The original task note framed members as client-owned. Every model inherited it; the audit verified mechanics, not domain semantics. Only the domain owner caught it. Eval oracles are not a substitute for owner review.
The best build still needed four adversarial rounds (7 findings → 2 → 2 → 0) plus a domain-model correction before earning DEPLOY-SAFE. "Won the bake-off" and "safe for production" are different bars.
Research + builds cost ~$13 of OpenRouter credit — five times the per-round budget. Credit now sits at $22.6 of $25; top up before the next open-model round. Frontier legs ran on subscriptions.