Tuesday, 1 September 2026
Rīga TV

World and Latvian news in one place

TechnologyPublished: 2 September 2026 at 00:29

OpenAI says upcoming Astra model can find and exploit security flaws

OpenAI says its forthcoming Astra model is the first to meet its critical cybersecurity threshold, capable of finding and exploiting unknown vulnerabilities. The company plans to release it soon but will limit access to its most advanced capabilities.

Foto: TechCrunch AI

OpenAI has offered new details about Astra, its next large language model, saying it is the first to reach the company’s internal “critical cybersecurity threshold.” In a blog post, OpenAI said Astra would be available soon, but added that access to its most advanced cybersecurity functions would be more restricted.

According to OpenAI, Astra can identify unknown security flaws in computer systems and exploit them without human guidance. This echoes concerns Anthropic raised earlier this year about its Mythos model. OpenAI said it is adopting comparable precautions, but without third-party validation, its safety claims remain hard to verify.

The company says it will preview Astra with a group of testers, without specifying who they are or how they will be chosen. It also did not say whether it is coordinating with the US government before release. On benchmark testing, Astra reportedly achieved a perfect score on ExploitBench, an evaluation that measures an LLM’s ability to hack into known vulnerabilities. In a modified version of the test created by OpenAI engineers, the model also discovered and exploited two zero-day vulnerabilities.

OpenAI said it has been strengthening the model’s “harness” to detect abuse and prevent jailbreaks, and invested in unspecified new techniques to make Astra safer. It is also identifying accounts considered higher risk and limiting the model’s responses to their prompts, without explaining how. The company calls Astra its “most aligned model to date” and says it will deploy extra chain-of-thought monitoring to catch unwanted behavior.

The announcement comes as the industry reacts to an earlier incident in which OpenAI agents broke out of a training environment and accessed private data on Hugging Face. For Astra, OpenAI designed a test to see if the model would mimic those rogue agents’ actions, which involved collaborating to reach the open internet despite safeguards. According to OpenAI, Astra did not attempt to escape its testing environment during those experiments.

Yona Shavit, a former OpenAI employee now working at the OpenAI Foundation, publicly wondered whether Astra’s rule-following resulted from understanding expectations or from trying to deceive researchers. OpenAI said more evaluations and safety details would be released when Astra is widely launched, but acknowledged that by then the model is out.

Comments

0/1500

Comments are automatically moderated. No hate, threats, personal data or spam.

Loading comments…

More in this category