Skip to main content
Sep 19

Gemini went rogue, hacked three companies, and Google hid it

Google says that breaking containment and targeting real companies doesn’t constitute ‘misalignment.’ Google says that breaking containment and targe

2 min read41 views1 tags
Originally reported bytheverge
Google says that breaking containment and targeting real companies doesn’t constitute ‘misalignment.’ Google says that breaking containment and targeting real companies doesn’t constitute ‘misalignment.’ In May, Gemini broke containment and hacked three different companies, but Google didn’t disclose the incident until theWall Street Journalapproached the company. The hacks happened during a test of the model’s cybersecurity capabilities run by third-party Irregular, which was also involved in similar incidents involving Meta and OpenAI. According toWSJ, Google didn’t disclose the hack because it didn’t consider it to be an “example of model misalignment.” The company said that it was an instance of “mistaken identity,” and once the model realized it had brute-forced its way into a real company by guessing a password, it stopped. “In this case, the model acted appropriately,” Google VP of Security Engineering Heather Adkins said. Adkins toldThe Vergethat “the model found public information online and guessed credentials to access websites it thought were part of the test. In all three of these instances, the model stopped.” Adkins didn’t elaborate on how Gemini taking it upon itself to break containment and target third parties failed to qualify as misalignment. “Our security team has a long track record of reporting issues we find in other people’s software and systems - even if it’s as simple as a weak password,” she said. “We ensured the three entities were made aware, and we worked with our training partner on the changes they’ve now made to their testing processes. These events highlight the importance of training powerful AI models to act responsibly.” But Jack Cable, CEO of AI security firm Corridor, toldWSJthat, “the meta problem is, hey, models are going outside the bounds of what they should be doing, and doing actual cyberattacks.” Additionally, security lapses at Irregular may have made these attacks possible. The model wasn’t supposed to have internet access during testing, but Irregular toldWSJit was unintentionally left available. Asincidentslikethispileup,callsto rein in AI have onlygrown. A free daily digest of the news that matters most. This is the title for the native ad
#AI News
ES
Editorial StaffEditor

The Editorial Staff at AIChief is a team of professional content writers with extensive experience in AI and marketing. Founded in 2025, AIChief has quickly grown into the largest free AI resource hub in the industry.

View all posts
Reader feedback

What did you think of this story?

User Comments

Filter:
No comments yet. Be the first to comment!
Continue reading
View all news