brain/
conceptartificial-intelligence

Terminal-Bench harness swing

Notes

Vintage: 2026-04–08. Official Terminal-Bench 2.0 numbers are from the fetched leaderboard page as of the 2026-08-27 research pass (2026-08-27-best-agent-harnesses-for-programming). Codex CLI + GPT-5.5 dated 23 Apr 2026 on that page.

Terminal-Bench harness swing

One-line summary: Official Terminal-Bench 2.0 shows the same model swinging ~16–18 points across harnesses. Codex CLI + GPT-5.5 is the highest official vendor CLI on the fetched page (82.2%); Claude Code + Opus 4.6 sits at 58.0%. Specialized eval harnesses beat vendor defaults. Transfer to daily coding is open.

The insight

The April 2026 wiki already knew “same model can score differently in different agents” (ai-coding-benchmarks, MorphLLM). This source adds the official Harbor table and a same-model swing large enough that “best CLI” and “best model” come apart.

The chain

Harbor TB 2.0 holds the task set fixed; the same model swings ~16–18 points across vendor vs specialized harnesses; vendor-CLI rankings are therefore harness-contaminated; daily-coding transfer stays open.

Canonical: harness-choice-to-terminal-bench-swing.

Evidence

  • From 2026-08-27-best-agent-harnesses-for-programming: Official Terminal-Bench 2.0 is itself a Harbor eval harness (harbor run … -k 5, no timeout/resource overrides).
  • From the same source: Codex CLI + GPT-5.5 is rank 4 at 82.2%±2.2 (23 Apr 2026). Claude Code + Opus 4.6 sits at 58.0%.
  • From the same source: Claude Opus 4.6 is 58.0% in Claude Code, 74.7% in Terminus-KIRA, 75.3% in Capy, 76.4% in Stanford IRIS Meta-Harness.
  • From the same source: GPT-5.5 is 82.2% in Codex CLI vs 84.7% in NexAU-AHE. Top rows also include LemonHarness 84.5% and Capy + GPT-5.5 83.1%.
  • From the same source: OpenHands + Claude Opus 4.5 is 51.9%±2.9 — strong open platform, not the TB winner.
  • From the same source: LangChain claims a harness-only jump of Deep Agents from top-30 to top-5 on TB 2.0; the fetched leaderboard has Deep Agents + GPT-5.2-Codex at 66.5% (rank 27).
  • Older snapshot on ai-coding-benchmarks (April 2026 MorphLLM): Codex CLI ~77.3% and Claude Code ~65.4% on Terminal-Bench 2.0. Different date, different model pairing, different fetch. Do not silently overwrite; both vintages stand.

Terminal-Bench 4.0 on AA Index v4.3 (2026-09-07) — not a TB 2.0 rewrite

From 2026-09-07-artificial-analysis-intelligence-index-v4-3. AA’s Index constituent is TB v4.0, not this page’s Harbor TB 2.0 table. Do not overwrite the 82.2% / 58.0% rows.

  • From the same source (@ArtificialAnlys): 66 multi-step terminal tasks; three repeats; average pass@1; harness moved from Terminus 2 to mini-SWE-agent.
  • From the same source (same post): gpt-6-astra (max) 59.1%; claude-fable-5-1 (max with fallback) 52.0%; Claude Opus 5 (max) 49.0%; Astra 19 points ahead of GPT-5.6 Sol (max) 39.9%.
  • From the same source (TB4 page blurb): Astra (xhigh) 59.6%, Astra (max) 59.1%, Fable 5.1 (xhigh / default fallback) 55.1%. Same Astra (max); different Fable effort than the thread’s 52.0%. Prefer effort-labeled citations. Anthropic’s issuer Fable 5.1 55.8% on claude-fable-5-1 stays a third print.

Contradictions / tensions

  • Vendor default vs eval-optimized harness. Specialized harnesses beat Claude Code for Opus 4.6 by ~16–18 points. Anthropic and LangChain both say models are post-trained inside a vendor harness and can overfit (LangChain cites Codex apply_patch). Unresolved: how much of the TB gap is reward-hacking / timeout policy vs a real daily-coding gap. Harbor forbids timeout/resource overrides; that does not prove transfer to a laptop repo. See does-terminal-bench-harness-swing-transfer-to-daily-coding.
  • OpenHands “preferred eval harness” vs this table. openhands SDK docs claim preferred-eval status and SOTA on SWE-bench / SWT-bench / multi-SWE-bench. TB 2.0 has OpenHands well below Codex CLI. Those claims can both be true if they refer to different benches and dates; this pass did not fetch SWE-bench.org numbers.
  • April MorphLLM vs official TB 2.0 (this pass). Claude Code 65.4% → 58.0% and Codex CLI 77.3% → 82.2% are not a single time series. Different model pairings (Opus 4.5-era vs Opus 4.6; GPT-5.3-era vs GPT-5.5). Leave both cited.
  • AA TB 4.0 vs Harbor TB 2.0. Different task set (66 tasks), different harness (mini-SWE-agent vs Terminus / vendor CLIs). Do not treat 59.1% / 52.0% as a continuation of the 82.2% / 58.0% table. From 2026-09-07-artificial-analysis-intelligence-index-v4-3.
  • Fable TB4: AA thread 52.0% (max with fallback) vs AA page 55.1% (xhigh) vs Anthropic issuer 55.8%. Effort / issuer mismatch. Do not flatten. Transfer question stays open — does-terminal-bench-harness-swing-transfer-to-daily-coding.

Open questions

Sources

Related

Referenced by