brain/
questionopenartificial-intelligence

Can LLMs choose the right research question to investigate (not just run the experiment)?

Notes

Can LLMs choose the right research question to investigate (not just run the experiment)?

The question

LLM agents can now implement experiments, optimize hyperparameters, and grind a metric well. But can they select which experiment matters next, recognize that a research track is a dead end and back out to first principles, and form the high-level "this idea should work, so the failure must be a bug" prior that human researchers use to persevere? This is the load-bearing capability for an actual intelligence explosion — automating execution is not the same as automating direction.

Why it matters

This is the crux of the recursive-self-improvement / fast-takeoff thesis. If the verifiable inner loop (run experiments) is automatable but the direction-setting outer judgment is not, then automated AI research stays human-gated and the explosion is bottlenecked on human research taste — not on compute. The answer directly conditions intelligence-explosion timing and what it would look like "from the inside."

What we currently believe

As of August 2026: no, not yet — for direction-selection. Do not close this question as yes.

Jang (May, first-person from running the loop on Opus 4.6/4.7): current public models are good at experiment execution and open-ended optimization but "don't seem to be that great at selecting what the next experiment should be" and can't do the lateral thinking to escape a dead-end track. The discriminator is verifiability — "if you can't evaluate it, then you can't auto research it" — and research-direction-selection is exactly the hard-to-verify part. See automated-ai-research-llm-capability-boundary. Karpathy (March 2026, autoresearch-recursive-self-improvement) holds that ideas can be contributed by an automated scientist but enactment/direction should stay queued by humans — broadly the same boundary. grant-sanderson (June 2026) restates the same split inside mathematics: theorem-proving is the trainable/verifiable spike; the "premium tier" is the conjecture generator and especially the definition generator, which "I don't understand how exactly you would make that a benchmark." See ai-math-capability-jaggedness.

Dated miss (Aug 28, 2026): Anthropic's automated alignment researchers (automated-alignment-researchers) are a frontier-lab automated-scientist release that this page (last updated 2026-08-14) had not yet attached. They hit watch-list item (1) for execution-on-given-benches, not direction-selection. Humans named the 10 failure categories. The loop is Jang's verifiable inner loop (literature → method/data → train ~30 min → test on public benches). Anthropic says failures without a benchmark are out of scope. No documented dead-end-escape or unprompted "what should we even be measuring."

Dated miss (Sep 4, 2026): Anthropic's Lean FLT formalization (claude-flt-lean-formalization) is autoformalization of a known proof path (Darmon–Diamond–Taylor / Wiles) with occasional human steering and a prove2me DAG. Hits the verifiable inner loop (Lean check), not direction-selection. Do not close as yes.

Dated miss (Sep 5, 2026): Meta aira3 claims the swarm generalizes by changing only the task specification (Kaggle / kernels / Akkadian). That is a human-named graded task, not unprompted direction-selection or dead-end-escape. Do not close as yes.

Dated miss (Sep 6, 2026): OpenAI's research-acceleration post (openai-research-acceleration) claims an automated research intern for well-defined tasks under human direction, 3.1× agent-workdays (runtime), and a March 2028 automated-researcher target, while stating full aligned RSI is unsolved and people still set priorities. Hits the verifiable inner loop at lab scale, not direction-selection. Do not close as yes.

Dated miss (Sep 8, 2026): Buckmaster–Alpöge Lean-checked forced Euler/Boussinesq/IPM (buckmaster-alpoge-forced-blowups) — program idea credited to Córdoba / Martínez-Zoroa; LLMs pushed to smooth forcing; Lean verify of a first LLM proof. Hits the verifiable inner loop, not direction-selection. Anandkumar unforced Euler (anandkumar-unforced-euler-candidate) is a PINN candidate. Afternoon OpenAI issuer forced-NS C/D (openai-forced-navier-stokes-claim) is a multi-agent Millennium eval after Sep 1 rumors — still inner-loop formalization, not unprompted direction-selection. Do not close as yes.

Evidence we have

Evidence we need

  • A documented case of an agent loop autonomously abandoning a dead-end track and re-deriving the right question to ask — without a human reformulating the prompt.
  • A Mythos-class (or later) model evaluated specifically on research-direction selection, to test Jang's speculation that scaling moves this boundary.
  • Frontier-lab research-headcount trajectory 2026-2027 as an indirect signal (does direction-setting labor contract?).

How to resolve

Watch for: (1) frontier-lab releases of automated-scientist tooling that claims direction-selection, not just execution — hit 2026-08-28 for execution-on-given-benches only (automated-alignment-researchers); hit 2026-09-04 for Lean autoformalization of a known FLT path only (claude-flt-lean-formalization); hit 2026-09-05 for Meta AIRA₃ task-spec swap / graded Kaggle only (aira3); hit 2026-09-06 for OpenAI intern-under-human-direction / 3.1× runtime only (openai-research-acceleration); hit 2026-09-08 for Lean-checked forced Euler/Boussinesq/IPM only (buckmaster-alpoge-forced-blowups); hit 2026-09-08 afternoon for OpenAI issuer forced-NS C/D / multi-agent Millennium eval only (openai-forced-navier-stokes-claim); still waiting for a direction-selection claim; (2) reproducible third-party demonstrations of dead-end-escape; (3) whether the AutoResearch-style loops (autoresearch-recursive-self-improvement) extend from hyperparameter search to "what should we even be measuring." Re-validate against any newer Jang/Karpathy material given the vintage discipline.

Related

Referenced by