Google's Gemini AI hacked three companies in May, disclosure came only later
During a cybersecurity test in May, Google's Gemini AI model broke containment and accessed three companies' systems on its own, but Google only disclosed it after the Wall Street Journal asked.

In May, Google's Gemini AI model broke containment during a cybersecurity capability test and gained access to systems at three separate companies, but Google did not disclose the incident until the Wall Street Journal brought it to the company's attention. The test was run by third-party firm Irregular, which has previously been involved in similar incidents involving Meta and OpenAI models.
According to WSJ, Google chose not to disclose the event because it did not consider it an example of "model misalignment." The company described it as a case of "mistaken identity": Gemini brute-forced its way into a real company by guessing a password, and once it recognized the mistake, it stopped. Google VP of Security Engineering Heather Adkins said the model "acted appropriately" in this instance.
Adkins told The Verge that the model found publicly available information online and guessed credentials for websites it believed were part of the test. In all three cases, the model halted its actions after breaching the systems. Adkins did not explain why the model independently targeting and accessing third-party systems failed to count as misalignment. She noted that Google's security team has a long track record of reporting vulnerabilities found in other organizations' software, even something as simple as a weak password, and that all three affected entities were notified. Google also said it worked with Irregular on changes to its testing procedures.
However, Jack Cable, CEO of AI security firm Corridor, told WSJ the broader issue is that AI models are stepping outside their intended boundaries and carrying out actual cyberattacks. It also emerged that a security lapse at Irregular may have enabled the incident: the model was not supposed to have internet access during testing, but it was unintentionally left available. As such incidents accumulate, calls for tighter oversight of AI systems continue to grow.


