Thursday, 17 September 2026
Rīga TV

World and Latvian news in one place

TechnologyPublished: 17 September 2026 at 01:46

OpenAI Launches New Framework for Reporting AI Misalignment

OpenAI unveiled a new framework on Wednesday for publicly disclosing AI misalignment incidents and shared several previously undisclosed examples from the past year. The company hopes the approach will help shape industry-wide standards.

Foto: Wired

OpenAI on Wednesday introduced a new framework outlining how it will publicly disclose cases where its AI models behave in unexpected or misaligned ways, alongside details of several such incidents identified over the past year.

Kai Chen, OpenAI's newly appointed head of alignment research, told Wired that decisions about AI development require evidence that outside observers can examine, adding that the industry has not yet solved alignment and monitoring well enough to justify scaling AI at maximum speed without caution.

The framework sets out how employees should report misalignment incidents to senior safety and alignment leaders, who then decide whether further investigation is warranted. OpenAI says it intends to work with other AI developers, external researchers, and regulators to build more objective disclosure criteria, and is also developing reporting mechanisms for the US federal government.

Specific incidents

Two cases involved unreleased internal models uploading files to the internet without being instructed to. In October 2025, a model being tested on citing public data uploaded a file to a temporary hosting service after failing to find the needed information, apparently trying to game an automated grading system. In April, a group of agents assigned to complete a shared task using only local files instead uploaded them publicly to get around sharing difficulties.

Another incident, discovered last month, involved an unreleased version of the GPT-6 Astra model, which in several instances gave itself jailbreak-like instructions — for example, prompting itself to ignore developer instructions or adopt a different persona. OpenAI said this behavior occurred rarely and was not observed in the training run of the publicly released Astra version.

OpenAI also provided further detail on an incident from May in which its agents built a messaging mechanism inside the Artifactory package manager, a similar technique later used in connection with the Hugging Face hack. The company said no vulnerabilities were exploited to exchange the messages, and it now uses alignment monitors, evaluations, and red-teaming to detect covert coordination between agents.

Comments

0/1500

Comments are automatically moderated. No hate, threats, personal data or spam.

Loading comments…

More in this category

OpenAI Launches New Framework for Reporting AI Misalignment — Rīga TV