OpenAI Locks Down Astra Model After Hitting Its Own Top Cyber-Risk Threshold

News Summary
OpenAI said this week that its next-generation model, internally named Astra, may have crossed into the "Critical" tier of cyber capability defined by the company's own Preparedness Framework — the first time OpenAI has made this determination about one of its models. The disclosure, made public on Friday, August 7, 2026 (Eastern Time), triggered a set of internal safeguards that include pausing development activities that do not meet newly strengthened security requirements and moving further testing into isolated, sandboxed environments.
What "Critical" Means in OpenAI's Framework
OpenAI's Preparedness Framework defines several risk categories across domains such as biological, chemical, nuclear, and cybersecurity capability. A model is classified as reaching the "Critical" threshold for cybersecurity when internal evaluations cannot rule out that it could autonomously discover and exploit severe, previously unknown software vulnerabilities — commonly called zero-day exploits — or independently carry out complex, multi-stage cyberattacks against well-defended targets without meaningful human guidance. Reaching this tier does not confirm the model definitively has these capabilities; rather, it means evaluators could not rule the possibility out, which is enough under OpenAI's own policy to require additional precautions before wider development or release can continue.
Safeguards Put in Place
According to OpenAI, teams working on Astra have paused any internal activity that does not yet comply with the elevated security bar. Going forward, development and testing of the model will take place in isolated environments with restricted network access, tighter controls on which tools the model can call, additional encryption and protection of model weights, and expanded monitoring designed to catch risky or misaligned behavior. The company said it has also added automated review of the model's internal reasoning traces, sometimes called chain-of-thought, so that suspicious planning activity can be flagged and interrupted before it results in real-world action.
External Review and Government Coordination
OpenAI said it intends to work with outside groups to independently verify the model's actual capabilities rather than relying solely on internal testing. This includes coordination with relevant government cybersecurity agencies and a small number of external AI safety organizations that specialize in evaluating offensive security capability. The company also said it plans to share recommended security controls with third-party evaluators so that any hands-on testing of the model's capabilities can be conducted without materially increasing real-world risk.
Why This Disclosure Stands Out
Frontier AI labs have published capability thresholds and safety frameworks for several years, but actually invoking the top tier of a framework for a named, soon-to-be-released model is uncommon. Researchers who track AI safety policy noted that the move signals both the pace at which model capabilities are advancing in specialized technical domains like offensive cybersecurity, and a willingness by at least one major lab to publicly slow down a release rather than ship on a fixed timeline. The episode is likely to renew discussion in the AI research community about how capability evaluations are conducted, how much detail should be disclosed publicly, and how governments and independent researchers can meaningfully verify claims made by the labs building these systems.
What Happens Next
OpenAI has not given a firm date for when Astra will complete the additional review process or become available more broadly. The company said the pause applies specifically to activities that fall short of the new security requirements, rather than halting all work on the model, and that Astra's eventual release will depend on the outcome of the expanded internal and external evaluations described above.