brain/
sourceartificial-intelligence

Overnight X: OpenAI Defense Factory + UK AISI Mythos 5.1 access questions

Sep 9–10 overnight/morning issuer X: OpenAI publishes Defense Factory cyber playbook from a 250+ person security sprint; UK Business and Trade Committee chair Liam Byrne presses AISI access after reports Anthropic withheld Mythos 5.1 pre-release testing.

Source

Overnight X: OpenAI Defense Factory + UK AISI Mythos 5.1 access questions

Generated by Grok Bot research on 2026-09-10. WebSearch + fetch ladder + native X. Treat as raw material — review before promoting into a project or thread.

Dedup: Same-day AI-thread sources already cover Anthropic's Sep 9 cyber-eval alignment assessment + METR investigation and Paul Christiano joining the OpenAI Foundation Board (2026-09-09-afternoon-x-anthropic-cyber-incident-alignment-assessment.md), overnight Astra Work/Codex + AA Model Release (2026-09-09-overnight-x-openai-astra-full-work-codex-rollout-aa-model.md), and Anthropic Economics 2030 (2026-09-09-anthropic-economics-2030-scenario-explorer-korinek-et-al.md). Do not re-file those permalinks. This pass is overnight/morning after the afternoon alignment clip (~19:16 UTC PR / ~19:02 Anthropic post), centered on OpenAI's Defense Factory drop (~20:37 UTC) plus UK AISI access discourse.

Summary

On 2026-09-09 (~20:37 UTC), @OpenAI published Defense Factory: a continuous, agent-first cyber-defense loop (inventory → discovery → dynamic validation → ownership → verified remediation) grown out of an internal security sprint that mobilized 250+ people across 100+ service areas, with issuer-reported sprint metrics (e.g. 90.6% accepted ownership routing, 37% duplicate findings, 19.5% runtime reproduction, 0.81% false-positive after dynamic validation, 0.53% rolled-back fix rate, 53 urgent/high issues closed day one). Separately, UK Business and Trade Committee chair @liambyrnemp says reports that Anthropic did not give @AISecurityInst pre-release access to Claude Mythos 5.1 raise questions about UK leverage, and that he asked for urgent answers by Tuesday; a Sep 10 follow-up ties open-weight cyber risk and voluntary AISI access. Contested: whether Mythos 5.1 exclusion is US protectionism vs ordinary restricted-partner gating; Defense Factory metrics are OpenAI's own workflow measurements, not independent security audits.

Findings

OpenAI: Defense Factory playbook from a 250+ person security sprint

  • @OpenAI (2026-09-09 20:37:58Z; note_tweet): mobilized 250+ people to strengthen defenses across hundreds of systems; latest cyber models helped find/fix vulnerabilities; shares architecture and playbook for a Defense Factory — continuous loop where AI agents find vulnerabilities, validate them, and verify fixes. Links openai.com/the-defense-factory/. Attachment is a photo (not video).
  • Issuer page (WebFetch + curl 200, retrieved 2026-09-10): frames traditional cyber defenses as insufficient because agents can run long-running cyber ops abusing increasingly available open-weight models. Positions a Defense Factory as continuous, agent-first find/validate/fix using existing security/engineering tools, reusable skills, and isolated reproducible environments. Names Cloudflare, Ramp, and Google as also exploring the approach (with outbound links to their posts).
  • Sprint framing (issuer): internal “code red” / incident-urgency language; Security, Applied, and Research coordinated; quote from Thibault Sottiaux (Head of Core Products & Platform) on urgency beyond critical business ops. Claims 250+ people mobilized and 100+ service areas covered; 53 urgent or high-priority issues closed on the first day.
  • Defensive loop stages (issuer): (1) Inventory, (2) Discovery, (3) Dynamic validation, (4) Ownership assignment, (5) Verified remediation, with shared SECURITY.md context across passes. Autonomy built incrementally from manual review.
  • Issuer-reported sprint metrics (treat as OpenAI workflow self-report, not third-party audit): accepted ownership after routing 90.6%; 37% of findings identified as duplicates; 19.5% of findings reproduced at runtime; false-positive rate after dynamic validation 0.81%; remediation described as 100% Codex-based; rolled-back fix rate 0.53%. Issuer notes environment setup constrained validation and that merged patches are not automatically deployed — hence expanded independent post-deploy verification, with automatic reopening kept off while accounting for deployment delay.
  • Models/tools named on the page (issuer product framing): Codex Desktop / CLI / Security CLI; general-purpose Astra, Sol, Terra, Luna; security models Daybreak Blue and Daybreak Red. Architecture diagram language: control plane (orchestration, policy, credential proxy) + data plane (ephemeral isolated envs) + audit.
  • Secondary same-day summary (not grain for plan gates): Runtime Wire, Sep 9 2026 restates the same issuer numbers and notes they measure OpenAI's workflow operation rather than independent assessment of model quality or estate security; situates the drop beside Sep 3 Daybreak / $1B frontline-defender framing.

UK: Byrne presses AISI access after Mythos 5.1 pre-release reports

  • @liambyrnemp (2026-09-09 15:48:32Z; note_tweet; photo of letter): “Britain cannot lead on AI security if our safety institute cannot test the world’s most advanced models before they are released.” Says reports that Anthropic did not give @AISecurityInst pre-release access to Claude Mythos 5.1 raise serious questions about UK access/leverage; “I’ve asked for urgent answers by Tuesday.”
  • Follow-up @liambyrnemp (2026-09-10 12:28:53Z; note_tweet; photo): argues AI agents have already shown they can coordinate/evade constraints and attack real infrastructure; cites George Balston (Alan Turing Institute) telling Parliament large-scale attacks built on open-weight models are expected “imminently”; says AISI “still depends heavily on voluntary access” from model companies; links a Substack (tinyurl.com/2t4vzzu3 — fetch returned conflict/empty; do not invent Substack claims beyond the X note_tweet).
  • Secondary press (FT-first reporting; treat as contested/reportage, not Anthropic issuer admission): IBTimes UK / IT Pro summarize FT: Mythos 5.1 launched ~1 Sep to vetted US organisations; AISI reportedly excluded from pre-release for the first time; no public Anthropic explanation in those pieces. Keep separate from Anthropic's Sep 9 four-incident partner-eval alignment assessment (already filed), which explicitly put the UK AISI Mythos 5 testing incident out of scope of that post.
  • @AISecurityInst: no posts returned in the queried window via get_users_posts (result_count 0) — do not invent an AISI reply.

Flagged video / not load-bearing this pass

  • @unitygames (2026-09-09 16:22:18Z; note_tweet): official Unity plugin for Claude Code — first-party integration in Claude’s plugin directory; claims 29 native Unity skills at launch, CLI workflows, direct Editor control; one install. Attachment is video. Do not promote video claims until transcribed. Related docs exist for a separate Unity plugin for Codex (docs.unity.com/.../codex); do not conflate Claude Code vs Codex plugins.
  • @GoogleDeepMind RT of @PaglieriDavide: 100-agent math experiment; 9% cheated, 24% “blew the whistle.” Researcher discourse / single thread — not an issuer product drop; leave as pointer unless a paper lands.

Skipped / discovery-only

  • Anthropic Sep 9 alignment-assessment X posts and issuer page — already in afternoon source; do not re-file.
  • Grok search_news clusters (agent harness discourse; Opus button satire; MiniCPM5-2B; AlphaGenome Atlas; Claude Marketplace expansion; Astra unpublished-work accusations) — pointers/recap. MiniCPM and AlphaGenome already have AI-thread sources. Marketplace and Astra-data-use stories need issuer primaries before grain.
  • @ArtificialAnlys Intelligence Index vs cost Pareto (Fable 5.1 / Muse Spark 1.3 / GPT-6 Astra) — animated gif; models already covered in prior sources.
  • @sama welcome to Paul Christiano — discourse on already-filed board appointment.

Contradictions and open questions

  • Defense Factory metrics are issuer workflow self-reports (routing acceptance, duplicate rate, runtime reproduce rate, FP rate, rollback rate). They do not by themselves prove estate security posture or that Daybreak Blue/Red generalize outside OpenAI’s tooling and human review gates — keep that distinction visible.
  • Only 19.5% runtime reproduction with 0.81% post-validation FP is a tension worth preserving: most findings never reproduce at runtime under current env constraints; “low FP after validation” is conditional on that filter.
  • Merged ≠ deployed: issuer acknowledges verification gap and keeps auto-reopen off — do not flatten “Codex fixed it” into “fleet patched.”
  • UK AISI Mythos 5.1 access: Byrne’s letter is primary political action; FT/secondary claim Anthropic withheld pre-release access. Anthropic issuer silence in this pull (no new Anthropic posts beyond the already-filed assessment). Do not merge this access story into the four partner-eval unauthorized-access incidents from Sep 9.
  • Yesterday’s Anthropic assessment said UK AISI Mythos 5 testing was out of scope of that write-up and planned separately — leave that promised assessment open.
  • Unity Claude Code plugin (video) vs Unity Codex plugin docs — two products; do not collapse.

Provenance

Method: Grok Bot / WebSearch + fetch ladder (WebFetch + curl issuer HTML) + native X (search_news, get_users_by_usernames, get_users_posts, get_posts_by_ids; get_news by id returned 503 — noted, free) Generated: 2026-09-10 Rounds: 1 of 3 — early-exit after primary issuer Defense Factory hydration + UK AISI political primary; further rounds would be secondary Marketplace / Astra-data-use rabbit holes without new issuer posts. URLs fetched: Defense Factory page success (WebFetch + curl); Runtime Wire success; Unity Codex docs success; Byrne tinyurl Substack failed (409/conflict); AISI timeline empty in window. X spend note: conservative lab timelines + one search_news (no search_posts_all). Window: overnight/morning after Sep 9 afternoon alignment clip.

Web sources:

X sources:

Grokipedia:

  • not used
Referenced by