AIRA₃
Vintage: 2026-09. Primary evidence is Meta's official @AIatMeta X text follow-ups in 2026-09-07-weekend-x-ai-meta-aira3-kaggle-gold-openai-wiki (hydrated 2026-09-07). The thread opener (2096271545589190927) is video and was not transcribed — do not treat that permalink as grain. Gold / architecture / RSI claims are Meta-stated snapshots of that thread, not a fetched blog or independent board.
AIRA₃
One-line summary: Meta's next-generation autonomous AI research system, announced 5 Sep 2026, claiming Gold (8th of ~4,000) on a June NVIDIA live Kaggle that fine-tuned a 30B Nemotron — RSI language is Meta framing, not demonstrated RSI.
What it is
A Meta Superintelligence Labs research system: many long-running agents (model + coding-harness pairs) in isolated environments, coordinating asynchronously through a forum (hypotheses/findings) and a shared filesystem (artifacts), with no central controller. Grain is the hydrated issuer thread. No Meta blog HTML was retrieved in this pass.
Why it matters to this thread
Autonomous research loops, agent-harness pairing, and recursive-self-improvement framing are in-scope. This is the first dated Meta AIRA₃ claim the thread has as a citable issuer thread. Distinct from muse-voice-transcribe and muse-spark. Do not treat RSI language as achieved RSI. Do not promote secondary LinkedIn “first gold by an autonomous AI research system” — that superlative is not in the hydrated Meta posts.
Key facts (from 2026-09-07-weekend-x-ai-meta-aira3-kaggle-gold-openai-wiki)
Live Kaggle gold — Meta claim; opener is untranscribed video
- From 2026-09-07-weekend-x-ai-meta-aira3-kaggle-gold-openai-wiki (official @AIatMeta, 2026-09-05): Meta entered AIRA₃ in a June live Kaggle run by NVIDIA to fine-tune a 30B Nemotron toward better reasoning. All competitors had the same information and were graded on a private test set. Meta claims 8th of ~4,000 teams (Gold) and “outperforming human competitors who had access to the same frontier tools.” Thread opener is video — not transcribed; not grain. Architecture / ensemble / generalization live in the text follow-ups.
Swarm architecture (text)
- From the same source (text follow-up): no central controller; many long-running agents (model + coding harness pairs) in isolated environments, coordinating asynchronously via (1) a forum for hypotheses/findings and (2) a shared filesystem for artifacts. Meta says search strategies emerge as agents choose which discoveries to build on, and that compute compounds knowledge over time.
Live ensemble and post-hoc medals (text)
- From the same source (text follow-up): gold entry was GPT 5.5 (OpenCode) + Claude 4.8 (ClaudeCode). Post-hoc on the same private set: Muse Spark 1.2 (MuseCode) also gold-level; Muse Spark 1.1 (OpenCode) and GLM 5.2 (OpenCode) silver-level. See opencode, claude-code, muse-spark.
Generalization / RSI language (text) — framing, not demonstrated RSI
- From the same source (text follow-up): Meta claims the system generalizes by changing only the task specification; internal benchmark 27% latency reduction on production GPU kernels; gold-level on another Kaggle translating Akkadian tablets. Closing claim: “a system that compounds its own knowledge” aimed at accelerating AI research and unlocking recursive self-improvement. Treat RSI language as Meta’s framing. Canonical attach: autoresearch-recursive-self-improvement. Chain: aira3-swarm-to-claimed-rsi.
What this source does not establish
- Thread opener is untranscribed video. Do not cite 2096271545589190927 as spoken grain.
- No Meta blog HTML retrieved. Grain is the X thread.
- Gold is Meta’s claim on a June competition, announced Sep 5. External private grading is the strongest part of the claim; it is still issuer-stated.
- “First gold by an autonomous AI research system” appears in secondary LinkedIn commentary, not in the hydrated Meta posts. Do not promote that superlative as Meta fact.
- RSI is not demonstrated. Do not collapse into “RSI achieved.”
- Not a rewrite of muse-voice-transcribe, muse-spark, automated-alignment-researchers, or claude-flt-lean-formalization.
- No ticker, 8-K, or stock-market tag. Do not drag meta’s capex/silicon page.
Contradictions / tensions
- June run, September announce. Competition date and announcement date are different. Not flattened.
- RSI language vs demonstrated RSI. Same thread uses “unlocking recursive self-improvement” as a goal. Flag, do not collapse. See autoresearch-recursive-self-improvement.
- Direction-selection vs task-spec swap. “Generalizes by changing only the task specification” is Meta’s claim that the same system can be pointed at a new graded task. That is not a close of can-llms-choose-the-right-research-question (status stays open).
Open questions
- What would an independent writeup of the NVIDIA Kaggle private set show versus Meta’s 8th / ~4,000 Gold claim?
- Does “change only the task specification” include choosing which research question to ask, or only executing a human-named task? Do not close can-llms-choose-the-right-research-question.