Artificial Analysis Model Release pages
Vintage: 2026-09. Primary evidence is official @ArtificialAnlys X plus the fetched releases index and GPT-6 Astra (max) page in 2026-09-09-overnight-x-openai-astra-full-work-codex-rollout-aa-model (
method: grok-bot;x_video: false). AA product/UI + AA numbers — not OpenAI or Anthropic issuer model cards.
Artificial Analysis Model Release pages
One-line summary: AA’s Sep 9, 2026 Model Release pages — a product/UI to compare intelligence, cost, and speed across effort levels. Worked examples for gpt-6-astra and claude-fable-5-1. Distinct from the Index v4.3 methodology page artificial-analysis-intelligence-index.
The insight
Frontier models now ship with up to six effort levels (AA claim); same weights can yield different performance and cost. Release pages are the UI AA shipped to show that spread (Intelligence Index, cost per task, output speed, latency, side-by-side effort comparison, plus an Artificial Analysis Capability Index with named categories). They are not lab model cards.
Index 53 at Astra/Fable max was already filed from the Sep 7 v4.3 pass. This page records the new UI and the effort-range lines. Do not flatten the worked-example Intelligence range 46–53 into the releases-listing 45–53.
Evidence
Product/UI claim (2026-09-09)
- From 2026-09-09-overnight-x-openai-astra-full-work-codex-rollout-aa-model (@ArtificialAnlys, 2026-09-09 03:26:01Z; note_tweet): new Model Release pages to compare intelligence, cost, and speed across effort levels. Claims frontier models now ship with up to six effort levels; same weights can yield different performance/cost. Pages include Intelligence Index, cost per task, output speed, latency, side-by-side effort comparison, and Artificial Analysis Capability Index scores (Finance & Accounting, Legal, Healthcare & Medical, Strategy & Ops, Engineering, Economics).
- From the same source (@ArtificialAnlys + artificialanalysis.ai/models/releases, WebFetch 2026-09-09): issuer index for those pages.
Worked example vs releases listing (keep split)
- From the same source (same announce post; AA, not OpenAI/Anthropic): on Intelligence vs Time per Task, gpt-6-astra ranges 46–53 on the Artificial Analysis Intelligence Index and 1.6–8.2 minutes per task; claude-fable-5-1 spans 47–53 intelligence on 4.2–12.2 minutes per task.
- From the same source (releases listing): GPT-6 Astra (OpenAI · Sep 2026 · 6 variants; Intelligence 45–53, cost/task $0.82–$3.26, 1M context) and Claude Fable 5.1 (Anthropic · Sep 2026 · 5 variants; Intelligence 47–53, cost/task $2.37–$7.63, 1M context) among others. 45–53 vs 46–53 not flattened.
- From the same source (GPT-6 Astra (max)): AA 53 Intelligence Index at max; $3.26/task; $10/$50 per 1M in/out; 1M context; knowledge cutoff Apr 30 2026; released Sep 3 2026. Max 53 / $3.26 matches the already-filed v4.3 print on artificial-analysis-intelligence-index — not a new max. $10/$50 is AA repeating a list already on gpt-6-astra; do not re-litigate.
What this source does not establish
- Not a lab issuer card. Do not back-fill these ranges onto OpenAI or Anthropic docs.
- Capability Index categories are named, not ranked here. The clip does not list category scores. Do not invent ranks.
- Not a rewrite of Index v4.3 methodology (TB 2.1→4.0; AutomationBench-AA; private-test 40%→45%; FAQ 25% vs thread 30/20/30/20). That stays on artificial-analysis-intelligence-index.
- Not MiniCPM / GLM / Kimi re-scores.
- Photos/text only on this clip.
x_video: false.
Contradictions / tensions
- Worked-example Astra Intelligence 46–53 vs listing 45–53. Same clip, two AA surfaces. Not reconciled.
- AA $3.26/Index-task vs OpenAI $10/$50 per 1M. Different units; already noted on gpt-6-astra. Not collapsed.
Open questions
- Whether Capability Index category ranks on the releases listing should ever be promoted onto model entity pages (clip says treat with care; no ranks extracted this pass).
- How the six-vs-five variant counts map onto OpenAI
reasoning.effort(low/medium/high/xhigh/max) and Anthropic effort labels.