AGI definitions disagree — and the benchmarks are saturating
AGI definitions disagree — and the benchmarks are saturating
Vintage: 2026-05. Primary source recorded 2026-05-30 (Moonshots EP #260). Capability/timeline claims; treat as a mid-2026 snapshot and frame older-vs-newer practitioner views chronologically per ../../_meta/AI_CAPABILITY_TRACKING. Speaker attribution is source-attributed — AssemblyAI mangled the diarization (Wissner-Gross labeled "Salim Ismail - 1"); names below are inferred from content/show-notes, not from the label map.
One-line summary: By mid-2026 the "are we at AGI?" question has fractured into mutually-inconsistent definitions held by credible people at the same moment — Hassabis (2029, must derive relativity), "we've had AGI since ~2020" (generality via GPT-2/few-shot), and "we don't even have a definition of intelligence." Underneath the noise, the standard benchmarks are saturating (top models cluster within tenths of a point), so the live methodological question is what replaces them — most likely benchmarks built from unsolved scientific/engineering problems rather than known-answer tasks.
The insight
The disagreement is less about facts than about where each speaker plants the AGI flag, and the recording dates matter (per the vintage discipline). Two structural claims survive the definitional noise:
- Definitions all roughly overlap one ~10-year window — so with hindsight the "when" debate will look like hand-wringing over a single historic period (the Industrial-Revolution-had-no-single-year analogy).
- Benchmark saturation is real and is itself evidence — when a battery of evals all cluster near the ceiling, models will look indistinguishable, which is exactly the llm-as-commodity-thesis convergence seen from the evaluation side. The next benchmarks have to test unsolved problems (derive-new-physics, open scientific/engineering challenges), not recall known answers.
Evidence
- From 2026-06-01-podcast-moonshots-opus-4-8-beats-gpt-5-5-the-220b-openai-foundation (Diamandis relaying Demis Hassabis, May 2026): Hassabis tightened his AGI timeline to 2029 (aligning with Kurzweil), calling today's agents "a practice run" with "only a few years to prepare," and proposed the "Einstein test" — a model trained only on pre-1901 knowledge that independently derives special relativity; current systems can't.
- From 2026-06-01-podcast-moonshots-opus-4-8-beats-gpt-5-5-the-220b-openai-foundation (Wissner-Gross, source-attributed, May 2026): "I think we've arguably had some form of artificial general intelligence since … 2020." He reads generality as achieved with GPT-2 / few-shot learners ("generality through a combination of prompt engineering and compression of general human knowledge"), and calls Hassabis's 2029 framing "moving the goalpost … conveniently, maybe somewhat self-servingly" given Gemini "is not winning the race."
- From 2026-06-01-podcast-moonshots-opus-4-8-beats-gpt-5-5-the-220b-openai-foundation (Salim Ismail, source-attributed, May 2026): "we don't know what intelligence is … the IQ test measures two aspects … physical, spatial, emotional, spiritual intelligence … that's not even in the equation. So I call bullshit on this until we can have a clear definition and a test for what we mean by even intelligence." Prediction: "we're going to keep moving the goalposts … and then we're going to go, oh, AGI sentience."
- Benchmark saturation (Wissner-Gross, source-attributed, May 2026): "all of these benchmarks are saturating … it's only when we see radical new benchmarks that you'd expect to see more dispersion … we're in an era of superintelligence and it's very easy … to just saturate every obvious benchmark." He's pushing (at the DOE Genesis Mission event) for "a new set of benchmarks … able to capture scientific and engineering open unsolved problems." Datapoints: Opus 4.8 at 57.9% on Humanity's Last Exam (with tools), GDP-Val AA at 1890.
- The wiki's own internal proxy (Blundin, May 2026): the pod had set "50% on Humanity's Last Exam" as its informal AGI/self-improvement threshold — "and now we're at … 57.9%."
Kurzweil — the origin of the "definitional spread," and his two remaining blockers (May 2026)
The mid-2026 fracture is exactly what Kurzweil predicted in 1999, which makes him the authoritative anchor for this concept:
- ray-kurzweil in 2026-06-03-podcast-moonshots-why-agi-is-close-but-not-here-yet-ray-kurzweil-ep (May 2026): "I also figured that people would have slightly different definitions of AGI. So there'd be a three year period where people would predict AGI's here and that would start three years earlier, like 2026 … That was my prediction back in 1999." → The "everyone has a different definition" noise the concept describes is, on his account, a predicted feature of the 2026–2029 approach, not a sign of confusion.
- ray-kurzweil in 2026-06-03-podcast-moonshots-why-agi-is-close-but-not-here-yet-ray-kurzweil-ep (May 2026), the concrete blocker-checklist (his answer to "what's missing for AGI"): "I think we need two things … it doesn't really understand physics … Google has announced a project to do that that I think will take till 2029. And then robotics is behind large language models … We don't have robotics that can do that at any price [and it] needs to be made less expensive … I think that will come about 2029. We know what needs to be done." A falsifiable version of the Einstein-test concern — real physics-understanding + affordable general-purpose robotics, both ~2029.
- ray-kurzweil (May 2026) on the saturation-adjacent capability cliff: "Large language models have only been effective for the last six months. Like a year ago really wasn't usable" and "Large language models are better than doctors … about 50% higher … That was not true a year ago." His read on why benchmarks look near-solved: the underlying exponential is moving so fast that capability jumps within-year.
Why it matters to this thread
- It reframes the agi-timeline-decade-of-agents question: timelines diverge mostly because definitions diverge, so a productive wiki posture is to track capability on specific benchmarks (and their saturation) rather than adjudicate "is it AGI."
- The saturation point is the evaluation-side face of llm-as-commodity-thesis — converging eval scores and converging products are the same phenomenon.
- It sets up a falsifiable next milestone: a model passing an unsolved-problem benchmark (Einstein-test-style) would be the first eval that re-introduces dispersion and would genuinely move the "AGI" needle.
Contradictions / tensions
- Self-interest cuts every way. Hassabis's conservative 2029 (Wissner-Gross's read: buys DeepMind time while Gemini trails); Wissner-Gross's "already AGI" is its own strong prior; Salim's "it's all noise" sidesteps falsifiability.
- "Benchmarks saturate" vs "Einstein test unmet" can both be true: today's evals saturate because they test known-answer tasks, while the unsolved-problem frontier (derive new physics) is wide open. The saturation is of the current benchmark generation, not of capability.
- Source is a single optimism-forward podcast; treat the "superintelligence era" framing as the panel's prior, not consensus.
Related
- ray-kurzweil — predicted both AGI-2029 and the 3-year definitional spread back in 1999
- agi-timeline-decade-of-agents — the Karpathy-anchored timeline this debate sits beside
- llm-as-commodity-thesis — eval convergence ↔ product commoditization
- automated-ai-research-llm-capability-boundary — the "can it choose the next experiment / solve the unsolved" boundary that an unsolved-problem benchmark would test