Terminal-Bench harness swing
Vintage: 2026-04–08. Official Terminal-Bench 2.0 numbers are from the fetched leaderboard page as of the 2026-08-27 research pass (2026-08-27-best-agent-harnesses-for-programming). Codex CLI + GPT-5.5 dated 23 Apr 2026 on that page.
Terminal-Bench harness swing
One-line summary: Official Terminal-Bench 2.0 shows the same model swinging ~16–18 points across harnesses. Codex CLI + GPT-5.5 is the highest official vendor CLI on the fetched page (82.2%); Claude Code + Opus 4.6 sits at 58.0%. Specialized eval harnesses beat vendor defaults. Transfer to daily coding is open.
The insight
The April 2026 wiki already knew “same model can score differently in different agents” (ai-coding-benchmarks, MorphLLM). This source adds the official Harbor table and a same-model swing large enough that “best CLI” and “best model” come apart.
The chain
Harbor TB 2.0 holds the task set fixed; the same model swings ~16–18 points across vendor vs specialized harnesses; vendor-CLI rankings are therefore harness-contaminated; daily-coding transfer stays open.
Canonical: harness-choice-to-terminal-bench-swing.
Evidence
- From 2026-08-27-best-agent-harnesses-for-programming: Official Terminal-Bench 2.0 is itself a Harbor eval harness (
harbor run … -k 5, no timeout/resource overrides). - From the same source: Codex CLI + GPT-5.5 is rank 4 at 82.2%±2.2 (23 Apr 2026). Claude Code + Opus 4.6 sits at 58.0%.
- From the same source: Claude Opus 4.6 is 58.0% in Claude Code, 74.7% in Terminus-KIRA, 75.3% in Capy, 76.4% in Stanford IRIS Meta-Harness.
- From the same source: GPT-5.5 is 82.2% in Codex CLI vs 84.7% in NexAU-AHE. Top rows also include LemonHarness 84.5% and Capy + GPT-5.5 83.1%.
- From the same source: OpenHands + Claude Opus 4.5 is 51.9%±2.9 — strong open platform, not the TB winner.
- From the same source: LangChain claims a harness-only jump of Deep Agents from top-30 to top-5 on TB 2.0; the fetched leaderboard has Deep Agents + GPT-5.2-Codex at 66.5% (rank 27).
- Older snapshot on ai-coding-benchmarks (April 2026 MorphLLM): Codex CLI ~77.3% and Claude Code ~65.4% on Terminal-Bench 2.0. Different date, different model pairing, different fetch. Do not silently overwrite; both vintages stand.
Terminal-Bench 4.0 on AA Index v4.3 (2026-09-07) — not a TB 2.0 rewrite
From 2026-09-07-artificial-analysis-intelligence-index-v4-3. AA’s Index constituent is TB v4.0, not this page’s Harbor TB 2.0 table. Do not overwrite the 82.2% / 58.0% rows.
- From the same source (@ArtificialAnlys): 66 multi-step terminal tasks; three repeats; average pass@1; harness moved from Terminus 2 to mini-SWE-agent.
- From the same source (same post): gpt-6-astra (max) 59.1%; claude-fable-5-1 (max with fallback) 52.0%; Claude Opus 5 (max) 49.0%; Astra 19 points ahead of GPT-5.6 Sol (max) 39.9%.
- From the same source (TB4 page blurb): Astra (xhigh) 59.6%, Astra (max) 59.1%, Fable 5.1 (xhigh / default fallback) 55.1%. Same Astra (max); different Fable effort than the thread’s 52.0%. Prefer effort-labeled citations. Anthropic’s issuer Fable 5.1 55.8% on claude-fable-5-1 stays a third print.
Contradictions / tensions
- Vendor default vs eval-optimized harness. Specialized harnesses beat Claude Code for Opus 4.6 by ~16–18 points. Anthropic and LangChain both say models are post-trained inside a vendor harness and can overfit (LangChain cites Codex
apply_patch). Unresolved: how much of the TB gap is reward-hacking / timeout policy vs a real daily-coding gap. Harbor forbids timeout/resource overrides; that does not prove transfer to a laptop repo. See does-terminal-bench-harness-swing-transfer-to-daily-coding. - OpenHands “preferred eval harness” vs this table. openhands SDK docs claim preferred-eval status and SOTA on SWE-bench / SWT-bench / multi-SWE-bench. TB 2.0 has OpenHands well below Codex CLI. Those claims can both be true if they refer to different benches and dates; this pass did not fetch SWE-bench.org numbers.
- April MorphLLM vs official TB 2.0 (this pass). Claude Code 65.4% → 58.0% and Codex CLI 77.3% → 82.2% are not a single time series. Different model pairings (Opus 4.5-era vs Opus 4.6; GPT-5.3-era vs GPT-5.5). Leave both cited.
- AA TB 4.0 vs Harbor TB 2.0. Different task set (66 tasks), different harness (mini-SWE-agent vs Terminus / vendor CLIs). Do not treat 59.1% / 52.0% as a continuation of the 82.2% / 58.0% table. From 2026-09-07-artificial-analysis-intelligence-index-v4-3.
- Fable TB4: AA thread 52.0% (max with fallback) vs AA page 55.1% (xhigh) vs Anthropic issuer 55.8%. Effort / issuer mismatch. Do not flatten. Transfer question stays open — does-terminal-bench-harness-swing-transfer-to-daily-coding.
Open questions
- Does the TB swing transfer to daily coding? does-terminal-bench-harness-swing-transfer-to-daily-coding
- Can scaffolding quality be benchmarked independently of the model? Already open on ai-coding-benchmarks.
Sources
- 2026-08-27-best-agent-harnesses-for-programming
- 2026-09-07-artificial-analysis-intelligence-index-v4-3 — AA TB 4.0 (66 tasks; mini-SWE-agent); effort-labeled Astra/Fable prints. Not a TB 2.0 rewrite.