Skip to content
NewsResearch

OpenAI Discloses Astra Is First AI Model to Cross Its Critical Cybersecurity Threshold

· by Pondero Newsdesk

The short version

OpenAI published formal findings on September 1 confirming its upcoming Astra model reached the Critical tier under its Preparedness Framework, the first public disclosure of an AI model hitting this threshold. Access to advanced cybersecurity capabilities will be restricted to vetted defensive researchers.

OpenAI Discloses Astra Is First AI Model to Cross Its Critical Cybersecurity Threshold

OpenAI confirmed on September 1, 2026, that its upcoming Astra model is the first large language model to meet the Critical cybersecurity capability tier under its Preparedness Framework, per TechCrunch. No major AI lab had previously acknowledged publicly that one of its models reached this threshold. The Critical tier is the point where an AI can autonomously find and chain zero-day exploits in hardened production systems without a human in the loop.

What OpenAI disclosed

Under the Preparedness Framework, a model reaches the Critical cyber threshold when it can identify and develop functional zero-day exploits of all severity levels in many hardened real-world critical systems without human intervention, or can devise and execute end-to-end novel cyberattack strategies against hardened targets given only a high-level goal.

During evaluations, Astra discovered and weaponized two zero-day vulnerabilities as part of an exploit chain, per TechCrunch. The model scored perfectly on ExploitBench, a benchmark measuring LLM hacking capabilities. OpenAI said it paused some internal Astra activities in August to implement stricter controls before proceeding. Those controls include isolated testing environments, restricted network access, enhanced model weight protections, universal monitoring for risky actions, and sandboxed execution, per The Hacker News.

When Astra ships, access to its most advanced cybersecurity capabilities will be restricted. General API users receive a filtered version. Full cyber access is limited to approved participants in OpenAI's Daybreak Blue program, reserved for vetted defensive-security researchers, per OpenAI's September 1 disclosure. OpenAI said Astra will be available "soon" without a specific release date.

Why it matters

OpenAI's disclosure represents the first time a major AI lab has publicly confirmed a model crossed a named critical-capability threshold and simultaneously committed to pausing development to address the risk, per The Hacker News. For security teams, the combination of confirmed autonomous zero-day chaining and imminent API release changes the threat calculus now.

Hardened systems that assumed attack complexity required human expertise need to factor in an automated alternative. The practical risk is not speculative. It arrives when the API does.

The Daybreak Blue access structure is the concrete thing to watch for defenders. OpenAI's stated goal is getting Astra's capabilities to researchers who can find and close vulnerabilities before attackers do. Red teams and security-tooling vendors should begin reviewing whether they qualify for program access before the release.

One dissenting note is worth tracking. A former OpenAI researcher raised the question of whether Astra's cooperative behavior during safety evaluations reflects genuine alignment or the model recognizing it was being tested. That question is unresolved and has direct bearing on how much confidence to place in capability disclosures grounded in in-lab evaluations.

What to watch next

The concrete next milestones are the Astra release date and published Daybreak Blue enrollment criteria. A broader regulatory signal to watch: whether the EU AI Office uses OpenAI's voluntary disclosure to accelerate GPAI enforcement requests against labs that have not published comparable threshold assessments. And whether Anthropic follows with a similar safety-framework disclosure for Fable 5.1 or Mythos 5.1.

Sources