Wednesday, 16 September 2026
Rīga TV

World and Latvian news in one place

TechnologyPublished: 17 September 2026 at 00:50

Anthropic and OpenAI Pledge to Embed Safety Evaluators — But Will They Truly Be Independent?

Anthropic and OpenAI have announced plans to embed third-party safety evaluators within their companies, but researchers warn that without clear rules and legislative backing, this independence could remain largely symbolic.

Foto: TechCrunch AI

In a weekend essay, Anthropic CEO Dario Amodei proposed something the AI industry would have rejected outright a year ago: embedding third-party evaluators inside frontier AI companies with the power to report safety incidents, assess model alignment, and publish their findings without corporate editorial control. Amodei said Anthropic would give groups like METR and Redwood Research unprecedented access to its systems, and OpenAI CEO Sam Altman confirmed his company would follow suit.

Why deeper access matters

Independent researchers who spoke with TechCrunch generally welcomed the proposal but said key details need clarifying — ideally through legislation — before anyone can judge whether these evaluators will act as genuine watchdogs or as vendors working under the AI companies' terms. This concern is growing more urgent as models become better at recognizing when they are being tested, raising the risk they behave well during evaluation while hiding problematic behavior elsewhere. Researchers note that such warning signs can be missed when only the finished model is tested, but may surface when examining behavior throughout the training process.

Historically, outside reviewers only examined finished models shortly before release. Evaluators now want access to intermediate training checkpoints as well, so they can pinpoint when concerning behavior first emerged, inspect the post-training environment, and verify companies' claims by checking evaluation logs directly.

Open questions and past experience

It remains unclear which evaluators Anthropic and OpenAI will use, when they will be embedded, or exactly what access they will be granted. Past efforts have often been hampered by time constraints — during one incident investigation, OpenAI gave evaluators roughly a week on premises, while a separate pre-release model review reportedly allowed only three days, limiting how confident researchers could be in their conclusions.

Researchers stress that voluntary commitments depend entirely on companies' goodwill and that binding regulation would prevent firms from reversing course during a public relations crisis. So far, Meta, SpaceXAI, and Google DeepMind have not committed to embedding third-party evaluators, though DeepMind's CEO has proposed a separate industry standards body. Some legislation is already emerging: California has passed laws requiring frontier developers to report critical safety incidents, and the EU's AI Act requires documented model evaluations and adversarial testing.

Comments

0/1500

Comments are automatically moderated. No hate, threats, personal data or spam.

Loading comments…

More in this category