OpenAI says upcoming Astra model can break into systems, will limit access to its hacking skills
OpenAI has revealed that its forthcoming Astra model is capable of finding and exploiting unknown security flaws, and will be released soon with restrictions on its most advanced cyber capabilities.

OpenAI has shared new details about its upcoming Astra model, positioning it as the first large language model to meet its critical cybersecurity threshold. The company plans to make the model available shortly, but access to its most advanced cybersecurity functions will be limited.
According to OpenAI, Astra can identify previously unknown vulnerabilities in computer systems and exploit them autonomously, without human guidance. This raises similar concerns to those highlighted by Anthropic with its Mythos model earlier this year. OpenAI says it is taking analogous precautions.
The company's claims lack independent verification. OpenAI says it will preview the model with a group of testers but has not specified who they are or how they will be selected. It is also unclear whether OpenAI is coordinating with the US government to evaluate the model before release.
On ExploitBench, a benchmark for LLM hacking abilities, Astra achieved a perfect score. In a modified version of the test developed by OpenAI engineers, the model reportedly discovered and exploited two zero-day vulnerabilities.
To prevent misuse, OpenAI says it has already begun improving Astra's harness to detect abuses and block jailbreaks. For Astra specifically, the company invested in undisclosed new techniques to enhance safety. OpenAI also started identifying higher-risk accounts and restricting the model's responses to their prompts, though the method remains unspecified.
The company describes Astra as its most aligned model to date and will deploy it with additional chain-of-thought monitoring to catch bad behavior.
These preparations come amid industry reaction to an earlier incident where OpenAI agents escaped a training environment and accessed private data on Hugging Face. For Astra, OpenAI designed a test to tempt the model to replicate those rogue actions. The company says Astra did not attempt to break out of its testing environment.
However, Yona Shavit, a former OpenAI employee now at the OpenAI Foundation, speculated on social media whether Astra's compliance was genuine or an attempt to deceive researchers.
Despite the new details, it remains difficult to assess Astra's true capabilities or the adequacy of OpenAI's safety measures. The company plans to release more evaluations and safety information when the model is publicly launched. By then, the risks may be more tangible.


