brain/
conceptartificial-intelligence

OpenAI Astra crossed Preparedness Framework "Critical" cyber

Notes

Vintage: 2026-09. Primary evidence is OpenAI's official post Path to Astra dated September 1, 2026, as hydrated in 2026-09-02-grok-com-ai-news-digest-2026-09-02-fable-5-1-astra-atlas-g20. Capability snapshot — not a generally available model card.

OpenAI Astra crossed Preparedness Framework "Critical" cyber

One-line summary: On September 1, 2026 OpenAI said Astra is the first model it is designating at the Preparedness Framework Critical cybersecurity threshold; ExploitBench 100% is known-vulnerability exploit development, and two zero-days were on a separate Internal Port set — not one flattened bench.

The insight

The Grok recap flattened two evals into "perfect ExploitBench + zero-days" as one bench and called Astra a "model suite." Issuer wording splits them and says "Astra" / "upcoming models." Grain: keep the split.

Evidence

  • From 2026-09-02-grok-com-ai-news-digest-2026-09-02-fable-5-1-astra-atlas-g20 (OpenAI Path to Astra, September 1, 2026): OpenAI now believes Astra meets the Critical cybersecurity capability threshold under its Preparedness Framework — "the first model we are designating at this level."
  • From the same source (issuer definition of Critical): either (1) identify and develop functional zero-day exploits of all severity levels in many hardened real-world critical systems without human intervention, or (2) devise and execute end-to-end novel strategies against hardened targets from a high-level goal.
  • From the same source: ExploitBench: Astra achieved a perfect score of 100% on the benchmark that evaluates exploit development from known vulnerabilities.
  • From the same source: because of contamination concerns, OpenAI built "ExploitBench - Internal Port (June–August 2026)" (20 high-severity V8 vulns disclosed more recently). On that set Astra discovered and used two zero-day vulnerabilities as part of an exploit chain; OpenAI is disclosing them to maintainers.
  • From the same source: those Astra results "reflect capabilities with Daybreak Blue access, not the default production configuration."
  • From the same source: Astra is not generally available. OpenAI "plan[s] to make Astra available soon," with advanced cyber initially limited to testers, then Daybreak Blue for defensive use.
  • From the same source: ExploitGym appears in the same post as a separate Hugging Face–informed honeypot / alignment eval, not the 100% score.
  • From the same source (native X): @OpenAI, 4:30 PM ET Sep 1.
  • From 2026-09-02-x-ai-overnight-fable-5-1-astra-critical-reward-seeker (official @OpenAI X + Path to Astra, same permalink / page): Astra "reaches the Critical cybersecurity threshold"; "the first OpenAI model designated Critical on cyber"; advanced cyber "initially be limited to testers / Daybreak Blue defensive access." OpenAI states "Astra was not involved in the Hugging Face incident but that incident learnings informed stronger safeguards." Exact public release timing remains "soon," not a dated GA claim. The page says Astra can find previously unknown flaws and develop exploits across many well-protected systems without step-by-step human guidance — do not reproduce attack procedures.

What this source does not establish

  • Recap "perfect ExploitBench + zero-days" as one bench is aggregator/recap copy, also flattened on the 7min chip. Issuer splits known-vuln ExploitBench vs Internal Port zero-days.
  • Not a "model suite." Issuer says Astra / upcoming models.
  • Not generally available and not default production — Daybreak Blue access.
  • This page does not reproduce attack procedures, exploit steps, or Internal Port vulnerability details. Cite the designation and the eval split only.
  • Not a re-file of openai-hugging-face-incident, openai-frontier-rl-pause, or openai-private-safety-processing. Adjacent safety pages; different grain.
  • No ticker, 8-K, or stock-market tag.
  • Do not rank against Gemini 3.8 Flash Cyber benches. From 2026-09-03-gemini-3-8-flash-and-3-8-flash-cyber-sep-2-2026: Google-reported CWE-Bench / internal 20-language / Chrome / CyberGym numbers are not shown to share a harness with ExploitBench or Internal Port. Comparability caveat only — not a re-file of Astra evals. Full treatment: gemini-3-8-flash.
  • Sep 3 product rollout is a different event. From 2026-09-03-gpt-6-astra-rolls-out-computer-use-flagship: official docs + X announce GPT-6 Astra (gpt-6-astra) rolling out to Trusted Access today, paid ChatGPT + API "in the coming days." Today's docs still point at Path to Astra for misalignment monitoring. That is not a rewrite of ExploitBench / Internal Port / Daybreak Blue on this page. Do not collapse "available soon / testers" into GA. Full treatment: gpt-6-astra.
  • Sep 6 issuer GPU reallocation. From 2026-09-07-x-overnight-openai-rsi-research-acceleration-data-jensen-agi (issuer research-acceleration page): August 7 — preliminary evidence Astra may have critical cyber capabilities under the Preparedness Framework → higher-security environment requirements; following week Astra-class GPU allocation −59.2%, other model classes +17.2% (~85% offset). Dated operational follow-through, not a rewrite of ExploitBench / Internal Port / Daybreak Blue. Full treatment: openai-research-acceleration.
  • Sep 9 Defense Factory tooling list. From 2026-09-10-overnight-x-openai-defense-factory-uk-aisi-mythos-51-access (issuer Defense Factory): names Daybreak Blue and Daybreak Red (plus Astra / Sol / Terra / Luna) as models/tools in the playbook. Not a rewrite of ExploitBench / Internal Port / Critical designation on this page. Does not close the Daybreak-Blue-for-defenders availability question. Full treatment: openai-defense-factory.

Contradictions / tensions

  • Grok recap vs issuer on Astra evals. Recap: "perfect score on ExploitBench and finding zero-days." Issuer: 100% on ExploitBench for known vulnerabilities; zero-days were on a separate Internal Port set. ExploitGym is yet another eval.

Open questions

Sources

Related

Referenced by