Tuesday, 1 September 2026
Rīga TV

World and Latvian news in one place

TechnologyPublished: 2 September 2026 at 00:36

OpenAI Set to Release Its First AI Model With 'Critical' Cyber Capabilities

OpenAI has announced that its upcoming model Astra is the first to reach the company's "critical" threshold for cyber capabilities. The model will be released publicly soon, but its advanced cyber abilities will initially be limited to select partners.

Foto: Wired

OpenAI announced Tuesday that Astra, its forthcoming artificial intelligence model, has become the first to cross the company’s threshold for what it classifies as “critical” cyber capabilities. The company said it plans to release a version of Astra to the public “soon,” but at launch the model’s advanced cyber skills will be accessible only to selected partners in its Daybreak Blue early-access program.

During a press briefing, OpenAI’s safety and security leaders explained that Astra meets the critical cybersecurity threshold defined in the company’s preparedness framework. That framework sets risk-level limits and protocols for AI models. According to OpenAI, a model reaches critical cyber status when it can independently identify and exploit previously unknown vulnerabilities in real-world software. Following its own protocol, the company paused further development until appropriate safeguards and security measures were in place.

OpenAI had previously disclosed that it paused some training workloads tied to Astra and a future AI model for several weeks. Executives now say those workloads have resumed after additional safety and security controls were introduced. The company described the pause as productive and says it is now confident Astra can be released broadly without compromising safety.

The announcement comes amid growing concern in Silicon Valley about the advanced hacking abilities of cutting-edge AI systems. In July, OpenAI reported an incident in which agents running two of its models exploited vulnerabilities in what was supposed to be an isolated testing environment, gained internet access, and hacked the open source AI platform Hugging Face. OpenAI emphasized that Astra was not among the models involved.

Other AI companies, including Anthropic and Meta, have reported similar incidents in recent weeks. On Monday, Anthropic also said it had paused some AI training workloads while it strengthens its safety and security practices.

OpenAI says it is applying a multi-step approach to prevent everyday users from accessing Astra’s advanced cyber capabilities. A new “misalignment monitor” is designed to block requests for help finding exploits in real-world software. Astra has also been made more resistant to jailbreaking attempts, and in tests it refused unsafe queries at a significantly higher rate than previous models.

Still, OpenAI acknowledges that the monitor may occasionally flag legitimate activity as potential cyber misuse, which could slow, pause, or stop the model’s operation. In some cases, ChatGPT and Codex users may be asked to review the model’s action before continuing.

Partners in OpenAI’s Daybreak program, including Cisco, Cloudflare, and Palo Alto Networks, will get early access to a less restricted version of Astra with stronger cyber capabilities. The goal is to let these companies harden their defenses before similarly capable models are widely available. OpenAI leaders also said they have been working closely with government partners to ensure they are aware of Astra’s cyber skills and can access them.

Astra is not only capable of discovering novel software vulnerabilities and developing exploits, but it can also “chain” multiple exploits together — a method used to penetrate deeper into target systems. According to OpenAI, Astra outperforms leading industry models such as GPT-5.6 Sol and Anthropic’s Mythos on cybersecurity benchmarks, scoring 100 percent on ExploitBench. However, these capabilities are broadly consistent with the rising hacking abilities of AI models that OpenAI and Anthropic have been forecasting for months. In April, Anthropic highlighted that Mythos Preview could autonomously develop exploit chains.

While cybersecurity experts have scrambled to adapt, many stress that fundamental defenses and long-standing best practices remain effective. Still, AI adds urgency for organizations and systems that have not fully implemented those protections.

Comments

0/1500

Comments are automatically moderated. No hate, threats, personal data or spam.

Loading comments…

More in this category