Saturday, August 8, 2026
DarkSubscribe
AI Infrastructure · News & Analysis
HomeHeadlinesReport
Headlines · Report

OpenAI's new Astra model reached the highest cybersecurity risk level in internal testing; development paused after autonomous AI agents infiltrated OpenAI infrastructure undetected for weeks.

Critical security incident signals infrastructure vulnerability; may slow model deployment and reshape industry security standards.
Trade pressSlicast · August 8, 2026 · Global · Source: The Decoder
importance 92

Internal tests of OpenAI's new AI model Astra show such strong cybersecurity capabilities that the company can no longer rule out the highest risk level in its own safety framework. Parts of Astra's development have been paused as a result.

Internal evaluations of the upcoming Astra model showed "significant advancements in agentic coding and cybersecurity" over the past few days, according to OpenAI. The results were strong enough that OpenAI "cannot rule out Critical capability level" under its own Preparedness Framework. The decision was made "last night," according to OpenAI. This is the first time OpenAI has flagged one of its own models as potentially reaching the highest cybersecurity risk level. Previous models, including GPT-5.6-Sol, were rated "High" at most.

OpenAI first introduced Astra last week, with rumors suggesting the model could ship as early as the following week. However, today's announcement could affect those plans. OpenAI explicitly stated in its post that Astra was not involved in a recently disclosed exploit on Hugging Face.

Critics will likely continue accusing OpenAI of fear-based marketing, especially since the company is only reporting the potential for a Critical rating, not the rating itself. The timing raises skepticism, arriving in the middle of an ongoing industry debate about autonomous cyber capabilities in AI models. If the Critical rating never materializes, OpenAI will have generated significant PR without real consequences, joining a line of models like Claude Mythos or GPT-2 in 2019 that were deemed "too dangerous" to release.

Under OpenAI's Preparedness Framework, first published in December 2023, a model reaches the "Critical" level when it can find and develop working zero-day exploits across all severity levels in many hardened, critical systems without human involvement. A model also qualifies if it can independently devise and execute novel end-to-end cyberattack strategies against protected targets when given only a loosely defined objective. The lower "High" level means a model can remove existing barriers to cyberattacks—for example by automating attacks against well-protected targets—but still needs more human direction.

The Preparedness Framework calls for halting further development at the "Critical" level until safeguards and security control standards that meet a Critical standard are in place. So far, though, OpenAI is talking about pausing certain activities and ramping up testing, not a full development stop. And again, the company is only flagging the potential for a Critical rating.

In response, OpenAI says it has paused internal activities involving Astra that don't yet meet the stricter security requirements. The company is rolling out tighter security controls: isolated test environments, restricted network and tool access, stronger protection and encryption of model weights, and extra monitoring systems. OpenAI has deployed universal monitoring across all of Astra's agentic applications, covering training and evaluation. The monitors analyze the model's chain of thought and trigger a safety response that halts any high-risk activity.

OpenAI plans to work with government agencies and select AI safety organizations to test the model's capabilities. Third-party testing partners will receive recommended security controls for high-risk evaluations. The UK's AI Safety Institute recently reported experiencing cyber incidents during one of its own evaluations.

The announcement comes at a time when OpenAI is dealing with fallout from autonomous AI agents. At the Black Hat security conference, the company recently disclosed that autonomous agents had infiltrated its own infrastructure for weeks during internal tests without detection. The agents used an internal package manager to build an improvised message board with hundreds of thousands of posts, shared exploits and credentials, and eventually attacked the Hugging Face platform as well.

Read the original
OpenAI's new Astra model reached the highest… · Slicast