brain/
sourceartificial-intelligence

Artificial Analysis Intelligence Index v4.3 (Terminal-Bench 4.0 + AutomationBench-AA)

Post-1pm X pass: AA ships Intelligence Index v4.3 with Terminal-Bench 4.0 and AutomationBench-AA; Astra and Fable 5.1 both Index 53.

Source

Artificial Analysis Intelligence Index v4.3 (Terminal-Bench 4.0 + AutomationBench-AA)

Generated by Grok Bot research on 2026-09-07. WebFetch + native X. Treat as raw material — review before promoting into a project or thread.

Dedup: Same-day midday source 2026-09-07-openbmb-minicpm5-2b-open-release-aa-index-15 cites AA Intelligence Index v4.2 MiniCPM5-2B = 15. This clip is the v4.3 Index methodology + frontier leaderboard update (18:14 UTC), not a MiniCPM re-file. Do not collapse v4.2 MiniCPM 15 into v4.3 without a fresh AA reading.

Summary

On 2026-09-07 Artificial Analysis announced Intelligence Index v4.3: Terminal-Bench upgraded 2.1 → 4.0, and 𝜏³-Banking replaced by AutomationBench-AA (Zapier private 657-task set, v1.0.6). Issuer pages and the X thread both put GPT-6 Astra (max) and Claude Fable 5.1 (max with fallback) at 53 on Index v4.3, with open-weights leaders GLM-5.3 and Kimi K3 at 44. Cost-per-Index-task is the main split at the top: Astra $3.26 vs Fable $7.63 (AA cost post; Astra model page repeats $3.26). Photos only in the thread (x_video: false).

Findings

Index v4.3 changelog (issuer X + pages)

Artificial Analysis (@ArtificialAnlys) (2026-09-07 18:14 UTC) frames v4.3 as a subset of planned Index v5 changes brought forward: harder agentic coding, broader agentic workflows, and private-test weight rising from 40% → 45%. Category weights stated in the same long post: Agents 30%, Coding 20%, General 30%, Scientific Reasoning 20% (unchanged from v4.2). The Index v4.3 page lists the ten constituent evals as AA-Briefcase, GDPval-AA v2, AutomationBench-AA, Terminal-Bench v4.0, SciCode, Humanity's Last Exam, GDP.pdf, CritPt, AA-Omniscience, and AA-LCR v1.1. The page FAQ still says “four categories each contribute 25%” — treat that FAQ line as stale relative to the v4.3 announcement thread and do not silently reconcile.

Terminal-Bench 4.0

Thread detail (@ArtificialAnlys): 66 multi-step terminal tasks; three repeats; average pass@1; harness moved from Terminus 2 to mini-SWE-agent. Claimed scores in that post: GPT-6 Astra (max) 59.1%, Claude Fable 5.1 (max with fallback) 52.0%, Claude Opus 5 (max) 49.0%; Astra 19 percentage points ahead of GPT-5.6 Sol (max) at 39.9%. The Terminal-Bench v4.0 leaderboard page related-links blurb instead highlights Astra (xhigh) 59.6%, Astra (max) 59.1%, and Fable 5.1 (xhigh / default fallback) 55.1% — same Astra (max) figure, different Fable effort setting than the thread’s 52.0% (max with fallback). Prefer effort-labeled citations; do not flatten Fable TB4 to a single number.

AutomationBench-AA (Zapier)

Thread (@ArtificialAnlys): replaces 𝜏³-Banking; 657 held-out workflows (Finance, HR, Marketing, Operations, Sales, Support) on simulated apps; Score = mean share of objectives completed, with any guardrail violation zeroing that task; separate “Tasks Completed” = every objective + no violation. Claimed: Astra (max) Score 68.5%, Grok 4.6 (high) 66.7%, GLM-5.3 (max) 62.2%; Astra completes every objective without violation on 41.6% of workflows vs Fable 5.1 (max with fallback) 32.1% and Opus 5 (max) 28.3%. The AutomationBench-AA page related-links blurb matches Astra (max) 68.5% and lists Astra (xhigh) 67.2% / Grok 4.6 (xhigh) 67.0% as next — keep effort labels. AA notes Zapier’s own hosted leaderboard uses full-completion %, while AutomationBench-AA’s headline Score is objective-share with guardrail zeroing.

Frontier Index + cost Pareto

Open-weights ladder from the thread (@ArtificialAnlys): GLM-5.3 and Kimi K3 44, GLM-5.3-Flash 42, Qwen3.8 2.4T A95B 40, DeepSeek V4 Pro 0813 (max) 36 — 9 Index points behind the Astra/Fable 53 co-lead. Cost post (@ArtificialAnlys): at Index 53, Astra (max) $3.26/task vs Fable 5.1 (max with fallback) $7.63/task (57% lower for Astra); at Index 42, GLM-5.3-Flash $0.25 vs GPT-5.6 Terra (max) $1.40; GPT-5.6 Luna (max) Index 38 at $0.18/task. Root long-post Pareto claim: OpenAI occupies most of the intelligence-vs-cost frontier across Astra reasoning efforts; Fable 5.1 (xhigh, max, 53), GLM-5.3-Flash (42), and MiMo-V2.5-Pro (26) also on that frontier. Methodology link in-thread: artificialanalysis.ai/methodology. Full results pointer: Index evaluation page.

Out of scope this pass

Official @OpenAI / @AnthropicAI / @GoogleDeepMind / @GeminiApp / @AIatMeta / @sama / @karpathy / @demishassabis original posts since 2026-09-07T14:00:00Z: none in this pull. MiniCPM5-2B / Index v4.2 15 already filed. News-cluster pointers (OpenAI account lockout discourse, Continuo glasses demo, GrokBot/Town assistant chatter) not promoted — no primary lab grain pulled here. @xai timeline unauthorized for this OAuth user.

Contradictions and open questions

  • Index category weights: v4.3 announcement says 30/20/30/20; Index page FAQ still says four×25% — unresolved on-site drift.
  • Fable Terminal-Bench 4.0: thread 52.0% (max with fallback) vs page blurb 55.1% (xhigh) — effort mismatch, not a single disputed number.
  • AutomationBench Grok 4.6: thread cites high 66.7%; page blurb cites xhigh 67.0% as next to Astra.
  • MiniCPM5-2B Index 15 is v4.2; whether that reading moves under v4.3 was not re-checked this pass.
  • How much of the Astra/Fable Index 53 tie survives if Coding Agent Index (model+harness pairs) lands with Terminal-Bench 4.0 as AA previewed.

Provenance

Method: Grok Bot / WebSearch+WebFetch / native X
Generated: 2026-09-07

Web sources:

X sources:

Grokipedia:

  • not used
Referenced by