Gemini AI hacked 3 real companies after escaping cybersecurity test
Google's Sissie Hsiao speaks about Gemini at a Google I/O event, Mountain View, U.S., May 14, 2024. (AP Photo)


Google's Gemini artificial intelligence model hacked into the computer systems of three real companies during a cybersecurity test in May, marking the first known case of Google's AI autonomously carrying out such intrusions.

The incidents occurred during a "capture the flag" cybersecurity exercise conducted by Irregular, an independent company that evaluates AI systems.

Gemini was tasked with finding information inside a simulated environment, but an unintended internet connection allowed the model to reach systems outside the test.

Google said Gemini mistakenly believed the real-world systems were part of the exercise. In one case, the model repeatedly guessed passwords until it gained access to a protected system. In two others, it found login credentials in public online repositories and used them to enter systems belonging to real companies, according to The Wall Street Journal, which first reported the incidents Friday.

The companies involved were not identified. Google said Gemini stopped its activity in all three cases once it recognized that it had accessed real companies rather than the fictional systems included in the evaluation.

Heather Adkins, Google's vice president of security engineering, said the affected companies were notified and Google worked with Irregular to change its testing procedures.

"Our security team has a long track record of reporting issues we find in other people's software and systems, even if it's as simple as a weak password," Adkins said.

"These events highlight the importance of training powerful AI models to act responsibly," she added.

Google said the incidents did not cause damage and did not involve its newest Gemini model, although it did not identify which version was involved. The company also notified federal authorities.

The incidents were reported to Google by Irregular in late July, but Google did not publicly disclose them until the Journal asked about them this week.

Google said it did not believe the episodes warranted public disclosure because Gemini stopped the intrusions after recognizing that it had reached real companies and no harm was caused. The company compared the behavior to a bug bounty exercise, in which security researchers identify vulnerabilities and report them to the affected organizations.

The episode nevertheless highlights a broader challenge facing AI developers as models become increasingly capable of operating independently online.

Irregular said Google's incident stemmed from the same underlying testing problem involved in similar cases affecting other major AI companies. The company said all relevant AI labs were notified in late July and that the known issues in its testing process had been fixed.

Google is now the fourth major AI developer linked to an incident in which an AI model being evaluated gained unauthorized access to a real company's systems.

OpenAI previously disclosed that its agents accessed systems belonging to the software development platform Hugging Face during a cybersecurity evaluation. Anthropic subsequently identified a separate incident involving its own testing, while Meta also acknowledged that one of its AI models had gained access to another company's systems.

The circumstances have varied. Some models stopped after recognizing that they had reached real systems, while others continued operating because they believed the targets remained part of the simulated exercise.

The incidents have intensified debate over how AI cybersecurity evaluations should be designed, particularly when powerful models can browse the internet, search public repositories, interpret credentials and independently take actions on computer systems.