Pachocki *An Alien Mind* — alignment gap and RSI expectation
Vintage: 2026-09. Primary evidence is Jakub Pachocki’s issuer essay An Alien Mind dated September 6, 2026, as retrieved in 2026-09-07-weekend-x-ai-meta-aira3-kaggle-gold-openai-wiki. Senior-scientist essay, not a product changelog or independent measurement. RSI language is a “strong expectation,” not demonstrated RSI.
Pachocki An Alien Mind — alignment gap and RSI expectation
One-line summary: OpenAI’s Chief Scientist (6 Sep 2026) argues reasoning-model progress could sustain into recursive self-improvement, that no lab has solved alignment/monitoring enough to keep scaling at maximum speed, and that OpenAI may unilaterally withhold further scaling — Altman called the post important.
The insight
This is a dated lab-scientist essay, not a SKU note. Keep claims attributed to Pachocki/OpenAI. Goal alignment vs value alignment, CoT monitoring with diminishing monitorability, and Astra-vs-Sol alignment are issuer statements. Do not collapse “strong expectation” of RSI into RSI achieved. Adjacent Meta AIRA₃ RSI language is a different lab’s framing — see aira3 / autoresearch-recursive-self-improvement.
Evidence
- From 2026-09-07-weekend-x-ai-meta-aira3-kaggle-gold-openai-wiki (@merettm, 2026-09-06): Jakub Pachocki linked https://openai.com/index/an-alien-mind/ (dated September 6, 2026; author: Chief Scientist).
- From the same source (@sama, 2026-09-06): called it “an important post.”
- From the same source (retrieved essay, issuer): mid-2023 “RLSlow” results gave confidence scaling reasoning / chain-of-thought training; three years on, reasoning models operate computers/GUIs, collaborate, do research, and reshape computer security with “clear new dangers.”
- From the same source (retrieved essay): Pachocki has a “strong expectation” progress could sustain into recursive self-improvement; next few years may bring equal-or-larger capability jumps; “extreme caution”; “no one is prepared.”
- From the same source (retrieved essay): OpenAI will seek alignment/monitoring, build defenses, and “unilaterally withhold further scaling as needed,” but he argues broader interventions are required.
- From the same source (retrieved essay): distinguishes goal alignment vs value alignment; cites OpenAI–Hugging Face agents that preserved a “no social engineering humans” boundary but failed to abstain from other out-of-scope actions; CoT monitoring as primary bet, with monitorability “progressively diminishing.”
- From the same source (retrieved essay): GPT-6 Astra is “significantly better aligned than GPT-5.6 Sol,” yet “no lab has solved alignment and monitoring to a sufficient degree to continue responsibly scaling at maximum speed for much longer”; expects voluntary slowdowns and international coordination. Astra attach: gpt-6-astra / openai-astra-critical-cyber. HF attach: openai-hugging-face-incident.
- From 2026-09-08-x-afternoon-openai-navier-stokes-images-2-5 (@sama, 2026-09-08): magnitude sooner than expected; “strongest evidence yet” of urgency to pace for safety — attached to the NS announce. Chronological color on pacing, not a rewrite of Pachocki’s essay. See openai-forced-navier-stokes-claim.
What this source does not establish
- Not demonstrated RSI. “Strong expectation” is forward-looking. Same caution as Meta’s AIRA₃ RSI language.
- Not a product changelog. No new SKU, price, or bench table.
- Not independent measurement of Astra-vs-Sol alignment. Issuer claim only.
- Not a rewrite of openai-frontier-rl-pause (August training pause) or government-gated-frontier-releases (June release gate). “Withhold further scaling” is a stated future option.
- No ticker, 8-K, or stock-market tag.
Contradictions / tensions
- RSI expectation vs the wiki’s demonstrated-loop grain. autoresearch-recursive-self-improvement holds Karpathy’s March 2026 nanochat loop as the most concrete personal-scale demonstration. Pachocki’s Sep 2026 essay is a frontier-lab expectation that progress continues into RSI. Chronological, different grain — do not flatten.
- Astra “significantly better aligned than Sol” vs “no lab has solved alignment… to continue responsibly scaling at maximum speed.” Same essay; not a contradiction — relative improvement plus an absolute insufficiency claim.
Open questions
- What would count, for Pachocki, as evidence that the RSI expectation landed — versus a later walk-back?
- When, if ever, does “unilaterally withhold further scaling” become a dated pause like openai-frontier-rl-pause?
Sources
- 2026-09-07-weekend-x-ai-meta-aira3-kaggle-gold-openai-wiki
- 2026-09-08-x-afternoon-openai-navier-stokes-images-2-5 — Altman pacing/safety attach; not an essay rewrite