OpenAI Defense Factory playbook from a 250+ person security sprint
Vintage: 2026-09. Primary evidence is official @OpenAI (2026-09-09 20:37:58Z) plus the fetched Defense Factory page (curl 200, retrieved 2026-09-10) in 2026-09-10-overnight-x-openai-defense-factory-uk-aisi-mythos-51-access (
method: grok-bot;x_video: false). Issuer playbook + workflow self-report — not an independent security audit. Do not flatten merged patches into fleet-deployed. This page does not reproduce attack procedures.
OpenAI Defense Factory playbook from a 250+ person security sprint
One-line summary: On 9 Sep 2026 OpenAI published Defense Factory — a continuous, agent-first find/validate/fix loop grown out of an internal security sprint that (issuer) mobilized 250+ people across 100+ service areas; sprint metrics are OpenAI’s own workflow measurements.
The insight
This is a lab-official cyber-defense playbook, not a Preparedness Framework designation and not a model SKU. Distinct from openai-astra-critical-cyber (Sep 1 Critical cyber threshold / ExploitBench split) and from claude-security (Anthropic Mythos 5 scans). The loop is inventory → discovery → dynamic validation → ownership → verified remediation, with shared SECURITY.md context. Autonomy is described as incremental from manual review. Keep issuer metrics visible as self-report, not as proof of estate posture or that Daybreak Blue/Red generalize outside OpenAI’s tooling and human review gates.
Evidence
Issuer X (2026-09-09)
- From 2026-09-10-overnight-x-openai-defense-factory-uk-aisi-mythos-51-access (@OpenAI, 2026-09-09 20:37:58Z; note_tweet): mobilized 250+ people to strengthen defenses across hundreds of systems; latest cyber models helped find/fix vulnerabilities; shares architecture and playbook for a Defense Factory — continuous loop where AI agents find vulnerabilities, validate them, and verify fixes. Links openai.com/the-defense-factory/. Attachment is a photo (not video).
Issuer page (WebFetch + curl 200, retrieved 2026-09-10)
From the same source (The Defense Factory):
- Traditional cyber defenses framed as insufficient because agents can run long-running cyber ops abusing increasingly available open-weight models.
- Defense Factory positioned as continuous, agent-first find/validate/fix using existing security/engineering tools, reusable skills, and isolated reproducible environments.
- Names Cloudflare, Ramp, and Google as also exploring the approach (outbound links on the page — not fetched as grain here).
- Sprint framing: internal “code red” / incident-urgency language; Security, Applied, and Research coordinated; quote from Thibault Sottiaux (Head of Core Products & Platform) on urgency beyond critical business ops. Claims 250+ people mobilized and 100+ service areas covered; 53 urgent or high-priority issues closed on the first day.
- Loop stages (issuer): (1) Inventory, (2) Discovery, (3) Dynamic validation, (4) Ownership assignment, (5) Verified remediation, with shared SECURITY.md context across passes.
- Issuer-reported sprint metrics (workflow self-report, not third-party audit): accepted ownership after routing 90.6%; 37% of findings identified as duplicates; 19.5% of findings reproduced at runtime; false-positive rate after dynamic validation 0.81%; remediation described as 100% Codex-based; rolled-back fix rate 0.53%.
- Issuer notes environment setup constrained validation and that merged patches are not automatically deployed — expanded independent post-deploy verification; automatic reopening kept off while accounting for deployment delay.
- Models/tools named (issuer product framing): Codex Desktop / CLI / Security CLI; general-purpose Astra, Sol, Terra, Luna; security models Daybreak Blue and Daybreak Red. Architecture diagram language: control plane (orchestration, policy, credential proxy) + data plane (ephemeral isolated envs) + audit.
Secondary same-day summary (not grain)
- From the same source (Runtime Wire, Sep 9 2026): restates the same issuer numbers and notes they measure OpenAI’s workflow operation rather than independent assessment of model quality or estate security; situates the drop beside Sep 3 Daybreak / $1B frontline-defender framing. Recap-only. Do not use Runtime Wire numbers as an independent audit.
What this source does not establish
- Not an independent security audit. Routing / duplicate / runtime-reproduce / FP / rollback rates are OpenAI workflow self-reports.
- Not a rewrite of openai-astra-critical-cyber. Daybreak Blue/Red appear here as named tooling on the playbook page — not a new ExploitBench / Internal Port / Critical-threshold claim.
- Not a Codex CLI SKU change. “100% Codex-based” remediation is issuer workflow language on codex-cli, not a permissions / AGENTS.md / price rewrite.
- Merged ≠ deployed. Do not flatten “Codex fixed it” into “fleet patched.”
- This page does not reproduce attack procedures, exploit steps, or how to run the loop offensively.
- Cloudflare / Ramp / Google “also exploring” is issuer naming, not fetched third-party playbooks.
- Runtime Wire is secondary. Same-day restatement; not grain for plan gates.
- Did not re-file afternoon alignment-assessment / Christiano board / overnight Astra Work-Codex / AA Model Release permalinks.
x_video: false. Unity Claude Code plugin in the same clip is video and not load-bearing — see claude-code.
Contradictions / tensions
- 19.5% runtime reproduction vs 0.81% post-validation FP. Most findings never reproduce at runtime under current env constraints; “low FP after validation” is conditional on that filter. Keep both numbers; do not flatten to “low false positives.”
- Self-report vs estate posture. Metrics do not by themselves prove OpenAI’s estate is secure or that Daybreak Blue/Red generalize outside this tooling and human review.
- Merged vs deployed. Issuer acknowledges the verification gap and keeps auto-reopen off.
Open questions
- Do Daybreak Blue/Red (or Astra / Sol / Terra / Luna as named here) generalize as defensive find/validate/fix outside OpenAI’s sprint tooling?
- What share of the 53 day-one urgent/high closures were actually deployed, versus merged and waiting?
- Is the 19.5% runtime-reproduce rate an environment-setup artifact the issuer can lift, or a durable filter on what the loop can verify?