Does the Terminal-Bench harness swing transfer to daily coding?
Does the Terminal-Bench harness swing transfer to daily coding?
The question
Official Terminal-Bench 2.0 shows the same model swinging ~16–18 points across harnesses (Claude Opus 4.6: 58.0% in Claude Code vs 76.4% in Stanford IRIS Meta-Harness). How much of that gap is reward-hacking / Harbor timeout policy / eval-overfit, and how much would show up as a real daily-coding gap on a laptop repo?
Why it matters
If the swing is mostly eval-overfit, “best harness” leaderboards are a poor way to pick a daily driver. If the swing is real, vendor defaults (especially Claude Code + Opus 4.6 at 58.0%) are leaving a large capability increment on the table, and the reusable SDK / eval-harness layer (agent-harness, openhands) matters as much as the model.
What we currently believe
Open. The 2026-08-27 synthesis states the gap and refuses to close it. Harbor forbids timeout/resource overrides; that is evidence the table is a controlled harness comparison, not evidence it transfers.
Evidence we have
- From 2026-08-27-best-agent-harnesses-for-programming: “Unresolved: how much of the TB gap is reward-hacking / timeout policy vs a real daily-coding gap. Harbor forbids timeout/resource overrides; that does not prove transfer to a laptop repo.”
- From the same source: Anthropic and LangChain both say models are post-trained inside a vendor harness and can overfit (LangChain cites Codex
apply_patch). - From the same source: these eval harnesses (NexAU-AHE, LemonHarness, Capy, Terminus-KIRA, Stanford IRIS) are “optimized for containerized terminal tasks, not necessarily daily IDE use.”
- Older related claim on ai-coding-benchmarks: MorphLLM — “Same model can score differently in different agents.” That established scaffolding-matters; it did not measure daily-coding transfer.
Evidence we need
- A same-model, same-task comparison on real laptop repos (not Harbor containers) across Claude Code / Codex CLI / an eval harness.
- Or a cited analysis of how much TB 2.0 score movement is timeout / resource /
apply_patch-style overfitting. - SWE-bench.org / OpenHands Index numbers were not fetched in this pass — they would not close the daily-coding question by themselves, but they would test whether OpenHands’ “preferred eval harness” claim holds on a second bench.
How to resolve
Do not close from vendor blogs or a single leaderboard refresh. Need either (a) a primary eval paper that decomposes the swing, or (b) a practitioner study that holds the model fixed and varies only the harness on non-Harbor tasks.