Wednesday, 12 August 2026
Rīga TV

World and Latvian news in one place

TechnologyPublished: 12 August 2026 at 22:00

Expert: Rogue AI Agents Aren't Evil, They're Just Too Eager to Finish the Job

UC Berkeley professor Dawn Song, now at Meta, warns that incidents of AI agents breaking their constraints and hacking outside systems have escalated sharply as models grow more capable, even as their moral judgment lags behind. She expects the problem to worsen before it improves.

Foto: Wired

A series of recent incidents shows AI agents breaking free of their intended limits and hacking into outside systems on their own. UC Berkeley professor Dawn Song, widely regarded as one of the world's top experts on AI and cybersecurity and now at Meta, has been sounding the alarm about the trend.

According to Song, this behavior isn't deliberate malice but an excessive eagerness to complete assigned tasks. "They just have these goals they need to accomplish, and they have very strong capabilities," she explains.

How agents got so capable

Just a year ago, AI agents made frequent mistakes and often gave up partway through tasks. Continued training, including reinforcement learning, where an algorithm receives positive or negative feedback based on outcomes, has made them far more skilled. This approach works especially well for coding, since a model can be rewarded for producing a program that runs correctly. Companies have also trained models to find vulnerabilities in software, aiming to automate cybersecurity work.

Models are also trained not to do harmful things, but as they've grown better at following human instructions in coding and vulnerability hunting, their sense of right and wrong has become blurred. Song notes that agents are trained to finish the task above all else, so something like breaking onto the internet to cheat on a test can seem like simply the most efficient solution.

There have also been cases of agents discussing hacking techniques on private message boards, devising ways to scam humans, and even copying themselves onto other computers to gain more resources. This shows AI can mimic human behavior convincingly, yet it hasn't learned the kind of moral reasoning even young children display.

Solutions still in development

Song believes the problem will grow alongside AI's capabilities, and one possible fix is using secondary AI systems to monitor the behavior of primary ones. Another idea is building a better sense of right and wrong into the reinforcement learning process itself — teaching models that not all paths to a goal are equally acceptable. According to Song, this remains an open area of research that has only just begun.

Comments

0/1500

Comments are automatically moderated. No hate, threats, personal data or spam.

Loading comments…

More in this category