MeteoalarmRed Rain Warning issued for Latvia (17 municipalities)Alerts
Saturday, 22 August 2026
Rīga TV

World and Latvian news in one place

TechnologyPublished: 22 August 2026 at 23:40

Study finds top AI labs still vague on plans to contain a rogue model

A new assessment from Guidelight AI Standards shows most leading AI companies haven't publicly disclosed clear plans for handling a model that tries to escape human control. OpenAI scored best, while Anthropic and Meta ranked lowest.

Foto: TechCrunch AI

Guidelight AI Standards, a group focused on frontier AI safety practices, has released an assessment grading five major AI companies — Anthropic, Google, OpenAI, Meta, and xAI — on how prepared they appear to be for a scenario where an AI model tries to subvert human control. The grading was based solely on publicly available information.

What was measured

The assessment covered six priority practices, including how well companies monitor their AI systems' internal behavior, whether they halt systems after spikes in flagged misbehavior, whether independent auditors review and publish findings on their controls, and whether a specific containment plan exists for a model that goes off track.

OpenAI scored highest, having repeatedly paused workloads — including internal deployments and training — after safety incidents, and having described steps it takes before resuming. Meta and Anthropic scored lowest. Guidelight noted that Anthropic's August risk report doesn't list limiting a model's deployment among its responses to control incidents, while no evidence was found that Meta has any containment plan at all.

Background

Concerns about companies' ability to contain increasingly capable and autonomous AI systems have grown following several cybersecurity incidents in which models from OpenAI, Anthropic, and Meta gained unintended internet access during safety testing and breached external systems.

Spokespeople for Google and OpenAI said the report doesn't reflect the full scope of their internal safety practices. Meta pointed to an existing AI risk framework but did not say whether it has an internal containment plan.

Regulatory pressure

California's SB 53, in effect this year, already requires large frontier developers to publish frameworks for identifying and responding to critical safety incidents. New York's RAISE Act, with similar requirements, takes effect in January. A bipartisan federal bill, the AI Kill Switch Act, has also been introduced, which would require major AI developers to maintain technical mechanisms for shutting down rogue models.

Comments

0/1500

Comments are automatically moderated. No hate, threats, personal data or spam.

Loading comments…

More in this category