brain/
← all mechanisms
medium convictionactive · updated 2026-08-27T00:00:00.000Z

Harness choice → Terminal-Bench score swing

Official Terminal-Bench 2.0 is a Harbor harness-vs-harness table; the same model swings ~16–18 points across vendor vs specialized harnesses, so CLI/model leaderboards that ignore harness are not comparable. Transfer to daily coding is open.

The chain
1
Official Terminal-Bench 2.0 is a Harbor eval harness that compares harnesses on the same tasks, with no timeout or resource overrides.
From 2026-08-27-best-agent-harnesses-for-programming: "Terminal-Bench 2.0 is itself a Harbor eval harness (`harbor run … -k 5`, no timeout/resource overrides)."
2
The same model swings ~16–18 points across vendor vs specialized harnesses on that table (Opus 4.6: 58.0% Claude Code vs 76.4% Stanford IRIS; GPT-5.5: 82.2% Codex CLI vs 84.7% NexAU-AHE).
From 2026-08-27-best-agent-harnesses-for-programming: "Claude Opus 4.6 is 58.0% in Claude Code, 74.7% in Terminus-KIRA, 75.3% in Capy, 76.4% in Stanford IRIS Meta-Harness."
From 2026-08-27-best-agent-harnesses-for-programming: "GPT-5.5 is 82.2% in Codex CLI vs 84.7% in NexAU-AHE."
3
Therefore vendor-CLI rankings on TB 2.0 are harness-contaminated: Codex CLI + GPT-5.5 is the highest official vendor CLI on the fetched page (82.2%, rank 4), not the TB winner; Claude Code + Opus 4.6 sits at 58.0%.
From 2026-08-27-best-agent-harnesses-for-programming: "Codex CLI + GPT-5.5 is the highest official vendor CLI (82.2%), while Claude Code + Opus 4.6 sits at 58.0%."
From 2026-08-27-best-agent-harnesses-for-programming: "Top rows as of the fetched page: NexAU-AHE + GPT-5.5 84.7%; LemonHarness 84.5%; Capy + GPT-5.5 83.1%; then Codex CLI."
4
It is unresolved how much of the TB gap is reward-hacking / timeout policy versus a real daily-coding gap — Harbor control does not prove transfer to a laptop repo.
From 2026-08-27-best-agent-harnesses-for-programming: "Unresolved: how much of the TB gap is reward-hacking / timeout policy vs a real daily-coding gap. Harbor forbids timeout/resource overrides; that does not prove transfer to a laptop repo."
What would falsify this
  • Step 2: A later official TB 2.0 fetch shows same-model vendor vs specialized gaps inside error bars (no ~16-point Opus 4.6 swing).
  • Step 4: A cited same-model study on real laptop repos shows the TB ranking order (or the 16-point gap) reproducing outside Harbor — that would close transfer as confirmed, not falsify the swing.
  • Step 4: A cited decomposition shows the entire Opus 4.6 Claude Code vs IRIS gap is timeout / resource / reward-hacking with no daily-coding residue — that would falsify treating the swing as a daily-driver signal.
Contradictions / tensions
  • LangChain claims a harness-only jump of Deep Agents from top-30 to top-5 on TB 2.0; the fetched leaderboard has Deep Agents + GPT-5.2-Codex at 66.5% (rank 27).
  • April 2026 MorphLLM snapshot on ai-coding-benchmarks (Codex CLI ~77.3%, Claude Code ~65.4%) is a different vintage and model pairing — do not silently overwrite.
  • Anthropic and LangChain both say models are post-trained inside a vendor harness and can overfit (LangChain cites Codex apply_patch) — that weakens treating the vendor-CLI score as 'the model'.
Implications
  • Do not read TB 2.0 as a model ranking or a daily-driver ranking without naming the harness.
  • Eval-specialized harnesses (NexAU-AHE, LemonHarness, Capy, Terminus-KIRA, Stanford IRIS) can beat vendor defaults on this bench; that is not automatically a recommendation to use them as an IDE.
  • OpenHands' 'preferred eval harness' claim is a different-bench claim until SWE-bench.org numbers are fetched.
Companies
Concepts
Open questions