Friday, 28 August 2026
Rīga TV

World and Latvian news in one place

TechnologyPublished: 28 August 2026 at 23:22

Anthropic researcher shows early AI system that improves model alignment on its own

Anthropic has published research on an automated system that reliably improves AI models' behavior on alignment benchmarks without hurting overall performance. The results are seen as an early step toward AI that can train and improve itself.

Foto: TechCrunch

A researcher in Anthropic's fellows program published a paper on Friday titled "Automated Researchers Can Reliably Mitigate Alignment Failures," describing an automated system designed to improve how AI models behave on alignment-related tests. The project was led by Anthropic fellow Chen Yueh-Han.

The system was tested against ten separate benchmarks, each measuring a specific type of misaligned behavior. The automated approach managed to improve scores on all ten benchmarks without reducing the models' overall performance.

How the process works

The system largely mirrors a traditional research workflow: it searches existing literature, proposes a training method, and then trains the model using that method for roughly 30 minutes, gradually raising benchmark scores across several iterations. Methods that prove effective are kept, while ones that don't work are discarded, allowing the process to run quickly and at scale.

The paper's authors describe the results as early evidence that automated alignment training could become practical in the near future. The work is being framed as a step toward recursive self-improvement — the idea that AI systems could eventually improve their own training processes.

Comparison with human researchers

The paper directly compares the performance of what it calls an Automated Alignment Researcher (AAR) to that of human experts. According to the authors, the best AAR method outperforms proposals from experienced human researchers within about six hours on average, and research directions guided by humans did not produce stronger results.

The paper also includes a cost comparison: running an AAR costs roughly $4 per hour in API inference, compared to about $150 per hour paid to human researchers.

The authors acknowledge limitations as well. The automated system's usefulness depends on how well the chosen benchmarks actually reflect real alignment goals, and significant work remains in building and maintaining those benchmarks, as well as expanding the body of literature the automated researchers draw from.

Comments

0/1500

Comments are automatically moderated. No hate, threats, personal data or spam.

Loading comments…

More in this category