brain/
conceptartificial-intelligence

Artificial Analysis Intelligence Index

Notes

Vintage: 2026-09. Primary evidence is one Grok Bot post-1pm X + fetch pass in 2026-09-07-artificial-analysis-intelligence-index-v4-3 (method: grok-bot; x_video: false). Grain is the fetched Index v4.3 page plus official @ArtificialAnlys X permalinks (18:14 UTC thread). Photos only. Do not collapse same-day 2026-09-07-openbmb-minicpm5-2b-open-release-aa-index-15 MiniCPM5-2B Index v4.2 = 15 into this v4.3 reading.

Artificial Analysis Intelligence Index

One-line summary: AA's composite Intelligence Index. v4.3 (7 Sep 2026) upgrades Terminal-Bench 2.1 β†’ 4.0, replaces 𝜏³-Banking with automationbench-aa, and raises private-test weight 40% β†’ 45%. Issuer thread + pages put gpt-6-astra (max) and claude-fable-5-1 (max with fallback) at 53.

The insight

v4.3 is framed as a subset of planned Index v5 changes brought forward: harder agentic coding, broader agentic workflows, and more private-test weight. The same long post states category weights Agents 30%, Coding 20%, General 30%, Scientific Reasoning 20% (unchanged from v4.2). The Index page FAQ still says four categories each contribute 25%. Treat that FAQ line as stale relative to the announcement thread. Do not silently reconcile.

This page is the Index methodology + v4.3 leaderboard. Model SKUs stay on their entity pages. MiniCPM5-2B 15 stays a v4.2 reading on minicpm5-2b β€” this pass did not re-check it under v4.3.

Evidence

v4.3 changelog (2026-09-07)

  • From 2026-09-07-artificial-analysis-intelligence-index-v4-3 (@ArtificialAnlys, 2026-09-07 18:14 UTC): Terminal-Bench upgraded 2.1 β†’ 4.0; 𝜏³-Banking replaced by AutomationBench-AA (Zapier private 657-task set, v1.0.6); private-test weight 40% β†’ 45%. Category weights in the same post: Agents 30% / Coding 20% / General 30% / Scientific Reasoning 20%.
  • From the same source (Index v4.3 page): ten constituent evals β€” AA-Briefcase, GDPval-AA v2, AutomationBench-AA, Terminal-Bench v4.0, SciCode, Humanity's Last Exam, GDP.pdf, CritPt, AA-Omniscience, AA-LCR v1.1.
  • From the same source (same Index page FAQ): β€œfour categories each contribute 25%.” Leave next to the thread’s 30/20/30/20. Unresolved on-site drift.

Frontier Index 53 + open-weights ladder

  • From the same source (issuer pages + announce thread): gpt-6-astra (max) and claude-fable-5-1 (max with fallback) both 53 on Index v4.3.
  • From the same source (open-weights post): GLM-5.3 and Kimi K3 44, glm-5-3-flash 42, Qwen3.8 2.4T A95B 40, DeepSeek V4 Pro 0813 (max) 36 β€” 9 Index points behind the Astra/Fable 53 co-lead.
  • From the same source (cost post): at Index 53, Astra (max) $3.26/task vs Fable 5.1 (max with fallback) $7.63/task (57% lower for Astra). At Index 42, GLM-5.3-Flash $0.25 vs GPT-5.6 Terra (max) $1.40. GPT-5.6 Luna (max) Index 38 at $0.18/task. Astra model page repeats $3.26.
  • From the same source (root long-post Pareto claim): OpenAI occupies most of the intelligence-vs-cost frontier across Astra reasoning efforts; Fable 5.1 (xhigh, max, 53), GLM-5.3-Flash (42), and MiMo-V2.5-Pro (26) also on that frontier. Methodology: artificialanalysis.ai/methodology.

Model Release pages (2026-09-09) β€” new AA UI, not a v4.3 rewrite

From 2026-09-09-overnight-x-openai-astra-full-work-codex-rollout-aa-model β€” AA product/UI + effort ranges. Canonical for the UI: artificial-analysis-model-release-pages. Not a methodology rewrite. Max 53 / $3.26 already on this page.

  • From the same source (@ArtificialAnlys, 2026-09-09 03:26:01Z): Model Release pages compare intelligence, cost, and speed across effort levels; up to six effort levels (AA claim). Worked example: gpt-6-astra 46–53 Index / 1.6–8.2 minutes per task vs claude-fable-5-1 47–53 / 4.2–12.2 minutes.
  • From the same source (releases index): Astra 6 variants, Intelligence 45–53, cost/task $0.82–$3.26; Fable 5 variants, Intelligence 47–53, cost/task $2.37–$7.63. 45–53 vs 46–53 not flattened.

What this source does not establish

  • Not a MiniCPM rewrite. minicpm5-2b Index 15 is v4.2. Whether that reading moves under v4.3 was not re-checked.
  • Not lab-issuer scores. Official @OpenAI / @AnthropicAI / @GoogleDeepMind / @GeminiApp / @AIatMeta / @sama / @karpathy / @demishassabis original posts since 2026-09-07T14:00:00Z: none in this pull.
  • Photos only. x_video: false. No video transcripts.
  • No new products. Zapier / mini-SWE-agent / Terminus 2 stay named artifacts. Did not create GLM-5.3 / Kimi K3 / Grok 4.6 / DeepSeek / Qwen3.8-A95B / Terra / Luna / MiMo pages.
  • Coding Agent Index (model+harness pairs) is previewed, not shipped in this pass.
  • Model Release pages are a product/UI, not a v4.3 methodology change. Capability Index category ranks were not extracted. From 2026-09-09-overnight-x-openai-astra-full-work-codex-rollout-aa-model.
  • Not a Speech-to-Speech Quality Index print. From 2026-09-15-google-gemini-3-8-live-3-8-live-extended-thinking: Google cites AA Speech-to-Speech Quality Index 82.6 / #1 for 3.8 Live Extended Thinking. That is a different AA instrument, Google-cited, AA page not hydrated. Do not fold 82.6 into Index 53.

Contradictions / tensions

  • Category weights: thread 30/20/30/20 vs FAQ fourΓ—25%. Unresolved on-site drift. Do not flatten.
  • Open-weights 44 vs closed 53 is a v4.3 Index gap, not a rewrite of the June GLM 5.2 / July Kimi K3 podcast evidence on chinese-open-weight-frontier-parity.
  • glm-5-3-flash Aug 27 β€œAA = 57” vs this pass’s Index v4.3 42 β€” different instruments / versions. Not reconciled. See that entity page.
  • AA Model Release worked-example Astra 46–53 vs listing 45–53. Same overnight clip. Not reconciled. From 2026-09-09-overnight-x-openai-astra-full-work-codex-rollout-aa-model.

Open questions

  • How much of the Astra/Fable Index 53 tie survives if Coding Agent Index (model+harness pairs) lands with Terminal-Bench 4.0 as AA previewed.
  • Whether MiniCPM5-2B’s v4.2 15 moves under v4.3.
  • Whether the Index page FAQ 25% line gets updated to match the v4.3 thread.

Sources

Related

Referenced by