medium convictionactive · updated 2026-08-14T00:00:00.000Z
Tight math verification → theorem-proving spike → century-scale definition lag → jagged AI-for-science progress
Math is AI's fastest field because contest proofs are trainable and tightly verifiable; the historically load-bearing work (conjectures, definitions, Langlands-style connections) has recognition loops that can take a century and will not fit current RLVR — so progress floods the verifiable layer while the conceptual layer stays human-gated (or later-AI).
The chain
1
Mathematics is a spike on AI's jagged frontier because IMO-style problems are trainable and tightly verifiable — and IMO gold is already old news as of June 2026.
grant-sanderson in 2026-06-30-podcast-dwarkesh-podcast-grant-sanderson-ai-and-the-future-of-math: "there's a spiky frontier to AI. Math is just right there in one of the spikes. But there's kind of a fractal nature to that spikiness, because when you zoom into the specific progress within math, you have some things that are a lot easier than others. So if we just think about imo, which is Old news at this point."
grant-sanderson in 2026-06-30-podcast-dwarkesh-podcast-grant-sanderson-ai-and-the-future-of-math: "I think the dirty secret with the IMO is that you really can train for a lot of them."
2
The premium-tier work — conjectures and especially definitions — cannot easily be made a benchmark, and historically can take ~100 years to be recognized even by human verifiers (Galois).
grant-sanderson in 2026-06-30-podcast-dwarkesh-podcast-grant-sanderson-ai-and-the-future-of-math: "how good mathematicians prove theorems. Good great mathematicians come up with conjectures and the greatest mathematicians come up with definitions. And that's more or less exactly your framing here on those two. We need the conjecture generator and then the definition generator. That's the premium tier mathematician. I don't understand how exactly you would make that a benchmark"
grant-sanderson in 2026-06-30-podcast-dwarkesh-podcast-grant-sanderson-ai-and-the-future-of-math: "What makes the Galois theory such an interesting example is you have literally this 100 year segment of, like, an idea that, like, flows through many different people's heads before it"
3
Even a correct AI proof can fail the actual goal (human understanding): proof ≠ explanation, and unsolved expository problems already exist in human math.
grant-sanderson in 2026-06-30-podcast-dwarkesh-podcast-grant-sanderson-ai-and-the-future-of-math: "everyone knows the idea of an unsolved research problem. I want to propose the idea of an unsolved expository problem, where, sure, we've proven it, but we don't really know why it's true."
grant-sanderson in 2026-06-30-podcast-dwarkesh-podcast-grant-sanderson-ai-and-the-future-of-math: "There is a difference between proof and explanation."
4
Near-term useful progress is therefore supercharged connections (Langlands character), still human-in-the-loop to score, while the remaining human role shifts toward curation rather than theorem-proving or even explanation.
grant-sanderson in 2026-06-30-podcast-dwarkesh-podcast-grant-sanderson-ai-and-the-future-of-math: "that's my guess on what most of the useful progress from these models will look like in the next five years is just really filling in that landscape of connections that you can draw."
grant-sanderson in 2026-06-30-podcast-dwarkesh-podcast-grant-sanderson-ai-and-the-future-of-math: "One interesting take that I've heard about what mathematicians will end up being is actually more analogous to art museum curators than anything else."
What would falsify this
- Step 1: A frontier lab reports IMO-style contest training no longer transfers to research-math tasks, i.e. the spike was contest-specific and dies.
- Step 2: A named new definition or axiom system generated by an AI is adopted into mainstream research math within ~2 years of publication (Galois-speed, not Galois-century).
- Step 4: Working mathematicians report that AI-proposed Langlands-style connections are being accepted without a human scoring step, for a year.
Contradictions / tensions
- Sanderson now suspects AIs will also be better explainers than most humans — which weakens step 3 as a durable human moat and pushes the remaining loop out to curation/motivation. Recorded, not treated as a break.
- Dwarkesh expects connection-making to be train-able via synthetic environments; Sanderson would be 'pretty surprised' if the next three years don't produce more lightning-bolt connections. Both expect step 4 to move; neither claims the definition generator has.
Implications
- Math is the leading indicator for other fields: wherever verification is cheap, AI floods; wherever recognition of a new *kind* of idea is slow, humans stay in the loop. See automated-ai-research-llm-capability-boundary.
- IMO / contest medals should not move agi-definitions-and-benchmark-saturation — Sanderson called that three years early.
- A model credited with a *definition* working mathematicians adopt would be the first eval that re-introduces dispersion on this chain.
Companies
Concepts
AI math capability is jagged — proofs spike; definitions and conjectures don'tWhat AI research LLMs can and can't automate (the capability boundary)AGI definitions disagree — and the benchmarks are saturatingCan LLMs choose the right research question to investigate (not just run the experiment)?
Open questions