Can LLMs choose the right research question to investigate (not just run the experiment)?
Can LLMs choose the right research question to investigate (not just run the experiment)?
The question
LLM agents can now implement experiments, optimize hyperparameters, and grind a metric well. But can they select which experiment matters next, recognize that a research track is a dead end and back out to first principles, and form the high-level "this idea should work, so the failure must be a bug" prior that human researchers use to persevere? This is the load-bearing capability for an actual intelligence explosion — automating execution is not the same as automating direction.
Why it matters
This is the crux of the recursive-self-improvement / fast-takeoff thesis. If the verifiable inner loop (run experiments) is automatable but the direction-setting outer judgment is not, then automated AI research stays human-gated and the explosion is bottlenecked on human research taste — not on compute. The answer directly conditions intelligence-explosion timing and what it would look like "from the inside."
What we currently believe
As of August 2026: no, not yet — for direction-selection. Do not close this question as yes.
Jang (May, first-person from running the loop on Opus 4.6/4.7): current public models are good at experiment execution and open-ended optimization but "don't seem to be that great at selecting what the next experiment should be" and can't do the lateral thinking to escape a dead-end track. The discriminator is verifiability — "if you can't evaluate it, then you can't auto research it" — and research-direction-selection is exactly the hard-to-verify part. See automated-ai-research-llm-capability-boundary. Karpathy (March 2026, autoresearch-recursive-self-improvement) holds that ideas can be contributed by an automated scientist but enactment/direction should stay queued by humans — broadly the same boundary. grant-sanderson (June 2026) restates the same split inside mathematics: theorem-proving is the trainable/verifiable spike; the "premium tier" is the conjecture generator and especially the definition generator, which "I don't understand how exactly you would make that a benchmark." See ai-math-capability-jaggedness.
Dated miss (Aug 28, 2026): Anthropic's automated alignment researchers (automated-alignment-researchers) are a frontier-lab automated-scientist release that this page (last updated 2026-08-14) had not yet attached. They hit watch-list item (1) for execution-on-given-benches, not direction-selection. Humans named the 10 failure categories. The loop is Jang's verifiable inner loop (literature → method/data → train ~30 min → test on public benches). Anthropic says failures without a benchmark are out of scope. No documented dead-end-escape or unprompted "what should we even be measuring."
Dated miss (Sep 4, 2026): Anthropic's Lean FLT formalization (claude-flt-lean-formalization) is autoformalization of a known proof path (Darmon–Diamond–Taylor / Wiles) with occasional human steering and a prove2me DAG. Hits the verifiable inner loop (Lean check), not direction-selection. Do not close as yes.
Dated miss (Sep 5, 2026): Meta aira3 claims the swarm generalizes by changing only the task specification (Kaggle / kernels / Akkadian). That is a human-named graded task, not unprompted direction-selection or dead-end-escape. Do not close as yes.
Dated miss (Sep 6, 2026): OpenAI's research-acceleration post (openai-research-acceleration) claims an automated research intern for well-defined tasks under human direction, 3.1× agent-workdays (runtime), and a March 2028 automated-researcher target, while stating full aligned RSI is unsolved and people still set priorities. Hits the verifiable inner loop at lab scale, not direction-selection. Do not close as yes.
Dated miss (Sep 8, 2026): Buckmaster–Alpöge Lean-checked forced Euler/Boussinesq/IPM (buckmaster-alpoge-forced-blowups) — program idea credited to Córdoba / Martínez-Zoroa; LLMs pushed to smooth forcing; Lean verify of a first LLM proof. Hits the verifiable inner loop, not direction-selection. Anandkumar unforced Euler (anandkumar-unforced-euler-candidate) is a PINN candidate. Afternoon OpenAI issuer forced-NS C/D (openai-forced-navier-stokes-claim) is a multi-agent Millennium eval after Sep 1 rumors — still inner-loop formalization, not unprompted direction-selection. Do not close as yes.
Evidence we have
- eric-jang in 2026-05-15-dwarkesh-podcast-eric-jang-building-alphago-from-scratch: models "don't seem to be that great at selecting what the next experiment should be in a given track ... I had to catch infra bugs myself. By prompting the right question to Claude."
- dwarkesh-patel in 2026-05-15-dwarkesh-podcast-eric-jang-building-alphago-from-scratch: the Ilya research-taste framing — a good researcher distinguishes "bug" from "wrong idea" via a strong high-level prior.
- Cross-reference: autoresearch-recursive-self-improvement — Karpathy's nanochat result (agents found tunings he missed) is the bullish counter-datapoint on the execution side; the direction side remains human in both accounts.
- grant-sanderson in 2026-06-30-podcast-dwarkesh-podcast-grant-sanderson-ai-and-the-future-of-math: "We need the conjecture generator and then the definition generator. That's the premium tier mathematician. I don't understand how exactly you would make that a benchmark." Galois as the century-scale recognition loop that doesn't look like RLVR.
- From 2026-08-31-anthropic-automated-alignment-researchers (Aug 28, 2026 Anthropic AAR; vintage 2026-08): Claude closed a substantial share of the safety gap on all 10 measured alignment-failure types via literature → method/data → train ~30 min on one H200 → test on public benches. Humans chose the 10 types (sycophancy, jailbreaks, prompt injection, power seeking, deception, hallucination, social bias, privacy violation, reward hacking, concealing uncertainty). Anthropic: failures without a benchmark are out of scope. Human-written initial directions did not improve AAR performance (long-form Sec. 5.1). See automated-alignment-researchers.
- From 2026-08-31-anthropic-automated-alignment-researchers (labeled count disagreement, keep both): blog "just over 2,000" training examples vs long-form "about 2,400" for Sonnet 5 → Opus 4.8. Prefer long-form if one count must be used. Harness announced open-source on the blog; no GitHub URL on the fetched page. Native X was 403; no permalink.
- From 2026-09-02-grok-com-ai-news-digest-2026-09-02-fable-5-1-astra-atlas-g20 (fetched abs only): arXiv:2609.01567 SAGE (agent guidance from imperfect VLM teachers) and arXiv:2609.01526 EvoSCM (closest fetched match to the recap's "causal model evolution for belief revision" / "LLM scientific law discovery" thread — not a paper titled "scientific law discovery"). Papers exist. Not a direction-selection claim and not a close of this question. Recap-named titles without a fetched abs stay recap-only. Status stays open. Do not close as yes.
- From 2026-09-04-anthropic-claude-formalizes-fermats-last-theorem-in-lean (Sep 4, 2026 Anthropic FLT Lean formalization; vintage 2026-09): Claude checked a known Darmon–Diamond–Taylor / Wiles path in Lean (11 days; Prove2Me DAG; occasional Peng steering). Hits the verifiable inner loop (autoformalization), not direction-selection, not a new conjecture, not dead-end-escape. See claude-flt-lean-formalization. Status stays open. Do not close as yes.
- From 2026-09-07-weekend-x-ai-meta-aira3-kaggle-gold-openai-wiki (Meta AIRA₃ text, 2026-09-05): “search strategies emerge as agents choose which discoveries to build on” and “generalizes by changing only the task specification.” Issuer framing on a specified Kaggle / kernel / Akkadian task. Not a documented case of abandoning a dead-end track without a human rewriting the question. See aira3. Status stays open. Do not close as yes.
- From 2026-09-07-x-overnight-openai-rsi-research-acceleration-data-jensen-agi (OpenAI issuer, September 6, 2026; vintage 2026-09): intern goal is well-defined tasks under human direction; “people still set priorities and decide scale/pause/deploy”; full aligned RSI unsolved. See openai-research-acceleration. Status stays open. Do not close as yes.
- From 2026-09-08-x-morning-buckmaster-alpoge-ai-fluid-proofs-openai-credit (Sep 8, 2026): Lean-checked forced blowups (Córdoba / Martínez-Zoroa idea → LLM smooth forcing / Euler). Inner-loop formalization, not unprompted direction-selection. See buckmaster-alpoge-forced-blowups. Status stays open. Do not close as yes.
- From 2026-09-08-x-afternoon-openai-navier-stokes-images-2-5 (Sep 8, 2026): OpenAI issuer forced-NS C/D after Sep 1 rumors / ~10k-agent eval. Hits the verifiable inner loop (issuer Lean), not direction-selection. See openai-forced-navier-stokes-claim. Status stays open. Do not close as yes.
Evidence we need
- A documented case of an agent loop autonomously abandoning a dead-end track and re-deriving the right question to ask — without a human reformulating the prompt.
- A Mythos-class (or later) model evaluated specifically on research-direction selection, to test Jang's speculation that scaling moves this boundary.
- Frontier-lab research-headcount trajectory 2026-2027 as an indirect signal (does direction-setting labor contract?).
How to resolve
Watch for: (1) frontier-lab releases of automated-scientist tooling that claims direction-selection, not just execution — hit 2026-08-28 for execution-on-given-benches only (automated-alignment-researchers); hit 2026-09-04 for Lean autoformalization of a known FLT path only (claude-flt-lean-formalization); hit 2026-09-05 for Meta AIRA₃ task-spec swap / graded Kaggle only (aira3); hit 2026-09-06 for OpenAI intern-under-human-direction / 3.1× runtime only (openai-research-acceleration); hit 2026-09-08 for Lean-checked forced Euler/Boussinesq/IPM only (buckmaster-alpoge-forced-blowups); hit 2026-09-08 afternoon for OpenAI issuer forced-NS C/D / multi-agent Millennium eval only (openai-forced-navier-stokes-claim); still waiting for a direction-selection claim; (2) reproducible third-party demonstrations of dead-end-escape; (3) whether the AutoResearch-style loops (autoresearch-recursive-self-improvement) extend from hyperparameter search to "what should we even be measuring." Re-validate against any newer Jang/Karpathy material given the vintage discipline.
Related
- automated-alignment-researchers
- agent-harness
- parametric-recall-bottleneck
- automated-ai-research-llm-capability-boundary
- autoresearch-recursive-self-improvement
- mcts-vs-llm-rl-credit-assignment
- agi-timeline-decade-of-agents
- eric-jang
- ai-math-capability-jaggedness
- grant-sanderson
- claude-flt-lean-formalization — Sep 4 Lean FLT; known path; do not close as yes
- openai-research-acceleration — Sep 6 OpenAI intern under human direction; do not close as yes
- prove2me
- aira3 — Sep 5 Meta swarm; task-spec / graded Kaggle; do not close as yes
- aira3-swarm-to-claimed-rsi
- buckmaster-alpoge-forced-blowups — Sep 8 Lean-checked forced blowups; do not close as yes
- anandkumar-unforced-euler-candidate — PINN candidate; do not close as yes
- openai-forced-navier-stokes-claim — Sep 8 afternoon issuer forced NS C/D; do not close as yes
- openai-unforced-euler — Sep 8 afternoon issuer unforced Euler; do not close as yes