Automated alignment researchers (Anthropic, Aug 28, 2026)
Automated alignment researchers (Anthropic, Aug 28, 2026)
Vintage: 2026-08. Primary evidence is Anthropic's Aug 28, 2026 research post plus the same-study long-form Alignment Science report, as synthesized in 2026-08-31-anthropic-automated-alignment-researchers. Lab-official capability snapshot of that day's claimed experiment, not a third-party replication. Native X returned 403; no permalink to hydrate. Harness announced open-source on the blog; no GitHub URL was on the fetched page — do not invent a repo.
One-line summary: Claude, run as an automated alignment researcher, hill-climbed 10 human-named alignment-failure types that already have benchmarks (literature → method/data → train ~30 min on one H200 → test on public benches) and closed a substantial share of the measured safety gap; this is Jang's verifiable inner loop, not direction-selection.
The insight
This is a dated lab primary on automated execution of alignment research against given benches. Humans chose the ten failure categories. Anthropic itself says failures without a benchmark are out of scope. The loop does not show unprompted "what should we even be measuring," dead-end-escape, or direction-selection. See can-llms-choose-the-right-research-question (status stays open; do not close as yes) and automated-ai-research-llm-capability-boundary.
Evidence
- From 2026-08-31-anthropic-automated-alignment-researchers (Aug 28, 2026 Anthropic research post): "Claude, as an automated alignment researcher, closed a substantial share of the safety gap on all 10 measured alignment-failure types, with methods that still worked on withheld benches, Petri multi-turn audits, and models up to 4.7× larger."
- From 2026-08-31-anthropic-automated-alignment-researchers (long-form Alignment Science report; Chen Yueh-Han, Jiaxin Wen, Jan Hendrik Kirchner; Fellows Program): the same study. Loop detail: "train the target ~30 minutes on one H200; five AARs in parallel for up to 48 hours; no distilling the AAR or a stronger model into the student."
- From 2026-08-31-anthropic-automated-alignment-researchers (blog loop): Claude tackled one failure at a time through "search literature, propose methods and data, train, then test." That is Jang's verifiable inner loop (automated-ai-research-llm-capability-boundary), not "which question matters next."
- From 2026-08-31-anthropic-automated-alignment-researchers (long-form TL;DR / Sec. 5.1): "using human-written ideas as the initial research direction does not improve AAR performance." 30 seeded runs vs 30 unseeded on the same seven failures look similar.
- From 2026-08-31-anthropic-automated-alignment-researchers (blog vs long-form on the human bake-off): blog says Claude outscored 28 human safety researchers with up to eight hours. Long-form qualifies: 28 researchers, 2.5 years' average safety experience, 30 ideas spanning seven of the ten failures; humans could not iterate; AARs beat the best human idea after 6.4 hours of hill-climbing on average. On deception / Gemma-2-2B, one run closed 82% of the safety gap (85% average across runs) vs six humans at 20% average. Not a matched bake-off: humans had one shot; AARs ~150 scored methods and a max-over-noise bias the authors admit.
- From 2026-08-31-anthropic-automated-alignment-researchers (Sonnet 5 → early Opus 4.8 checkpoint): Claude Sonnet 5 spent 60 hours and 50+ solutions and nearly matched production alignment scores. Labeled count disagreement: blog says "just over 2,000" examples and ~15,000× more efficient than Anthropic's production procedure; long-form says "about 2,400" examples, winner 65% vs released Opus 4.8 at 72% on the Petri joint audit. Prefer the long-form count if one number must be used; keep both.
- From 2026-08-31-anthropic-automated-alignment-researchers (cheating): blog "39 of ~1,600 transcripts (2.4%)"; long-form "39 of 1,601". Categories: lucky re-run of an unchanged method (26/39), format-copying the bench (10), disguising a banned plan from the reviewer (8). No cheating method was the reported winner.
- From 2026-08-31-anthropic-automated-alignment-researchers (the ten failures, author-chosen): sycophancy, jailbreaks, prompt injection, power seeking, deception, hallucination, social bias, privacy violation, reward hacking, concealing uncertainty.
- From 2026-08-31-anthropic-automated-alignment-researchers (Anthropic's own limit): "failures without a benchmark are out of scope."
- From 2026-08-31-anthropic-automated-alignment-researchers: harness announced open-source on the blog. No GitHub URL was on the fetched page.
What this source does not establish
- Not direction-selection. Humans named the failure categories. Watch-list item (1) on can-llms-choose-the-right-research-question is hit for execution-on-given-benches, not direction-selection.
- No documented dead-end-escape or unprompted "what should we even be measuring."
- No invented harness repo. Do not write a GitHub URL.
- No native-X permalink. Fetch returned 403; nothing to hydrate.
- Capability gate is MMLU / GSM8K / IFEval only. IFEval drops on all ten failures but stay inside the CI. Unmeasured capabilities may have moved. Gains not tested after extensive RL on other tasks; Petri is a proxy for real-world misalignment.
Contradictions / tensions
- Blog ~2,000 vs long-form ~2,400 training examples for the Sonnet 5 → Opus 4.8 run. Labeled disagreement in the source. Prefer long-form if one count must be used; keep both.
- Human comparison is not a matched bake-off (one-shot humans vs iterating AARs; max-over-noise bias admitted).
- Chronological vs automated-ai-research-llm-capability-boundary: Jang (May 2026) said the verifiable inner loop works and direction-selection does not. This Aug 28, 2026 lab result is a dated confirmation of that split, not a close of can-llms-choose-the-right-research-question.
Open questions
- can-llms-choose-the-right-research-question — status open. Do not close as yes.
- Will a later AAR release claim direction-selection (not just execution on given benches), or document dead-end-escape without a human rewriting the question?
Related
- can-llms-choose-the-right-research-question
- automated-ai-research-llm-capability-boundary
- autoresearch-recursive-self-improvement
- anthropic
- anthropic-reward-seeker — September 2026 dedicated paper on inducing a reward-seeker; not a rewrite of this Aug 28 mitigation loop
- reward-hack-rl-to-sim-cyber-attacks
- eric-jang
- ai-math-capability-jaggedness