AutomationBench-AA
Vintage: 2026-09. Primary evidence is 2026-09-07-artificial-analysis-intelligence-index-v4-3 (
method: grok-bot;x_video: false). Grain is the fetched AutomationBench-AA page plus @ArtificialAnlys. Photos only. Keep effort labels. Do not flatten Grok 4.6 high 66.7% into xhigh 67.0%.
AutomationBench-AA
One-line summary: Zapier-held-out workflow eval that replaces 𝜏³-Banking on artificial-analysis-intelligence-index v4.3. 657 simulated-app workflows. Headline Score = mean share of objectives completed, with any guardrail violation zeroing that task.
The insight
AA and Zapier’s own hosted leaderboard are not the same scalar. AutomationBench-AA’s headline Score is objective-share with guardrail zeroing. Zapier’s hosted board uses full-completion %. A separate “Tasks Completed” figure requires every objective and no violation.
Sep 3 gpt-6-astra rollout named AutomationBench with no scores. This page is the first dated AA print.
Evidence
- From 2026-09-07-artificial-analysis-intelligence-index-v4-3 (@ArtificialAnlys): replaces 𝜏³-Banking; 657 held-out workflows (Finance, HR, Marketing, Operations, Sales, Support) on simulated apps; Score = mean share of objectives completed, any guardrail violation zeros that task; “Tasks Completed” = every objective + no violation.
- From the same source (same post): gpt-6-astra (max) Score 68.5%; Grok 4.6 (high) 66.7%; GLM-5.3 (max) 62.2%. Astra completes every objective without violation on 41.6% of workflows vs claude-fable-5-1 (max with fallback) 32.1% and Claude Opus 5 (max) 28.3%.
- From the same source (AutomationBench-AA page related-links blurb): Astra (max) 68.5% matches the thread; next listed are Astra (xhigh) 67.2% and Grok 4.6 (xhigh) 67.0%. Keep effort labels.
- From the same source: Zapier’s hosted leaderboard uses full-completion %; AutomationBench-AA’s headline Score is objective-share with guardrail zeroing.
What this source does not establish
- Not a rewrite of the Sep 3 Astra named-board list. That pass had names, no scores.
- Not a Grok 4.6 product page. High vs xhigh scores stay here as effort-labeled prints.
- No video.
x_video: false. - Did not invent Zapier hosted-board percentages beyond AA’s statement that the scalar differs.
Contradictions / tensions
- Grok 4.6: thread high 66.7% vs page blurb xhigh 67.0%. Effort mismatch, not a single disputed number.
- Score vs Tasks Completed vs Zapier hosted % are three instruments. Do not collapse Astra 68.5% Score into 41.6% full-complete.
Open questions
- How Zapier’s hosted full-completion ranking would reorder the same models.
- Whether Coding Agent Index (model+harness) will sit beside this eval on a later Index.