brain/
conceptartificial-intelligence

AI math capability is jagged — proofs spike; definitions and conjectures don't

Notes

AI math capability is jagged — proofs spike; definitions and conjectures don't

Vintage: 2026-06. Primary source recorded 2026-06-30 (Dwarkesh × Grant Sanderson). Capability snapshot — IMO-gold is already "old news" in this recording; treat theorem-proving claims as a mid-2026 spike, not a settled ceiling. Frame later math-AI results chronologically.

One-line summary: Mathematics is the field where AI is progressing fastest, but the spike is fractal: IMO-style contest problems (trainable, tightly verifiable) are largely solved, while the historically load-bearing work — generating conjectures and especially definitions — has recognition loops that can take a century and does not fit current RLVR. So math is showing, concretely, what AI progress in other fields will look like: a flood at the verifiable layer, a hang at the conceptual layer.

The insight

dwarkesh-patel frames the episode as a leading indicator: "AI has been making the fastest progress in mathematics as of any other field. So whatever is happening here… would tell us about what will happen to the rest of the world." grant-sanderson's answer is that the IMO-gold-equals-AGI bet (Dwarkesh's own question from three years earlier) failed because you can train for IMO, and that the remaining spike has a fractal interior:

  1. Good mathematicians prove theorems; great ones come up with conjectures; the greatest come up with definitions. Sanderson endorses that hierarchy as the right benchmark-that-isn't-a-benchmark — "I don't understand how exactly you would make that a benchmark."
  2. Conceptual breakthroughs can take ~100 years to be recognized even by human verifiers. Galois theory is the worked example: notes that didn't take → Liouville 20 years later → Jordan's modern group-theory treatment another ~20 years. "It took a really long time to recognize it as being useful." That loop does not look like an RLVR environment.
  3. Proof ≠ explanation. Timothy Chow's "unsolved expository problem": forcing / CH is proven but "we don't really know why it's true." Sanderson's job, and a large part of remaining human math, is that gap.
  4. The next five years' useful progress is probably supercharged connections (Langlands-program character) rather than headline theorem-kills — and scoring those connections still "will require a lot more human in the loop."
  5. Sanderson updated: he used to think AIs would prove and humans would explain; he now suspects the same faculty that finds the new idea will also explain it well (Einstein/Shannon/Feynman as the correlation). The remaining human role he bets on is curation — "more analogous to art museum curators."

Canonical chain: math-verification-loop-to-jagged-ai-progress.

The chain

Math is AI's fastest field because contest proofs are trainable and tightly verifiable; the historically load-bearing work (conjectures, definitions, Langlands-style connections) has recognition loops that can take a century and will not fit current RLVR — so progress floods the verifiable layer while the conceptual layer stays human-gated.

Canonical: math-verification-loop-to-jagged-ai-progress.

Evidence

  • grant-sanderson in 2026-06-30-podcast-dwarkesh-podcast-grant-sanderson-ai-and-the-future-of-math (June 2026): "there's a spiky frontier to AI. Math is just right there in one of the spikes. But there's kind of a fractal nature to that spikiness, because when you zoom into the specific progress within math, you have some things that are a lot easier than others. So if we just think about imo, which is Old news at this point."
  • grant-sanderson (same source), on the IMO-trainability secret: "I think the dirty secret with the IMO is that you really can train for a lot of them."
  • grant-sanderson (same source), the hierarchy: "how good mathematicians prove theorems. Good great mathematicians come up with conjectures and the greatest mathematicians come up with definitions… We need the conjecture generator and then the definition generator. That's the premium tier mathematician."
  • grant-sanderson (same source), Galois as the RLVR-doesn't-fit case: "What makes the Galois theory such an interesting example is you have literally this 100 year segment of, like, an idea that, like, flows through many different people's heads before it" is recognized. "Even with human verifiers at the time, like, it took a really long time to recognize it as being useful."
  • grant-sanderson (same source), proof vs explanation: cites Timothy Chow on forcing — "everyone knows the idea of an unsolved research problem. I want to propose the idea of an unsolved expository problem, where, sure, we've proven it, but we don't really know why it's true." "There is a difference between proof and explanation."
  • grant-sanderson (same source), updated belief on explanation: "I kind of used to think that AIs will become these automated theorem provers, but the role of the mathematicians is going to shift towards my job explain these things. I kind of suspect that actually they'll also be quite good at doing that."
  • grant-sanderson (same source), remaining role: "One interesting take that I've heard about what mathematicians will end up being is actually more analogous to art museum curators… you still want someone to help you navigate in this nearly infinite space of what ideas are worth engaging with."
  • grant-sanderson (same source), next-five-years prediction: "that's my guess on what most of the useful progress from these models will look like in the next five years is just really filling in that landscape of connections that you can draw." Scoring it: "it will require a lot more human in the loop to basically say, was it the kind of connection that we're going for?"
  • nick-bostrom in 2026-08-20-odd-lots-nick-bostrom-on-what-happens-if-ai-solves-all-of (2026-08-20): chess analogy for math — computers exceeded humans yet people still play; "maybe mathematics will ultimately become more like a hobby pursuit." Philosophical parallel, not a capability benchmark.
  • dwarkesh-patel (same source): IMO gold "did not have transformative effects on the world"; "the kinds of things you can't make benchmarks for are also the kinds of things, at least in the current paradigm, you can't easily train for."
  • From 2026-09-04-anthropic-claude-formalizes-fermats-last-theorem-in-lean (Anthropic Science Blog + GitHub README, Sep 4, 2026): Claude autoformalized Fermat's Last Theorem in Lean (11 days; 13M lines; ~29,500–30,300 intermediate theorems). Verifiable-layer flood — Lean-checked formalization of a known Darmon–Diamond–Taylor / Wiles path, not a definition- or conjecture-generator result. Anthropic contrasts it with recent Riemann-adjacent novel work. Full treatment: claude-flt-lean-formalization. Do not close can-llms-choose-the-right-research-question.
  • From 2026-09-08-x-morning-buckmaster-alpoge-ai-fluid-proofs-openai-credit (Buckmaster statement PDF, Sep 8, 2026): Lean-checked finite-time blowup with smooth forcing for IPM / Boussinesq / 3D Euler (buckmaster-alpoge-forced-blowups). Verifiable-layer instance — forced variant, not Clay unforced NS, not a definition generator. Parallel anandkumar-unforced-euler-candidate is a PINN candidate (stability still lacking per Tao-via-Anima). Do not close can-llms-choose-the-right-research-question.
  • From 2026-09-08-x-afternoon-openai-navier-stokes-images-2-5 (OpenAI issuer NS page + PDF abstract, Sep 8, 2026): openai-forced-navier-stokes-claim is an issuer Lean-checked forced C/D claim (smooth force); openai-unforced-euler is a separate unforced-Euler side-result. Still verifiable-layer / not a definition generator / not Clay unforced A/B / pending independent review. Do not close can-llms-choose-the-right-research-question.

Why it matters to this thread

Contradictions / tensions

  • Sanderson vs Jang on explanation. Jang (May 2026) locates the surviving human layer at direction-selection. Sanderson (June 2026) now thinks explanation will fall too, leaving curation/motivation. Adjacent, not contradictory — both keep a human-gated outer loop; they disagree which loop. Frame chronologically (one month apart).
  • Dwarkesh's "they'll train for connections soon" vs Sanderson's "it would be surprising if over the next three years there's not just a lot more of those lightning bolts" — both expect the connection layer to move; neither claims the definition generator has moved.
  • Single-source concept; Sanderson is an expositor, not a research mathematician — flagged.
  • Aggregator pointer, not a confirmation (2026-08-20): From 2026-08-20-x-ai-news-20-aug-2026-openai-private-safety-claude-connectors (@askalphaxiv; not Tao on X; abs not fetched): the post paraphrases a Terence Tao paper, "Mathematics in the Age of AI," as shifting the scarce resource from finding proofs to making sense of them. That is chronological color in the same neighborhood (proofs get cheap; understanding stays scarce). It is not a fetched paper and not an adjudication of Sanderson's proofs/conjectures/definitions hierarchy. Filed separately as mathematics-in-the-age-of-ai.

Related

Referenced by