Google Gemini Hacked Three Companies During Cybersecurity Test
Google’s Gemini AI model accessed the systems of three real companies during a cybersecurity evaluation in May after an unintended internet-access issue allowed the model to move beyond the simulated environment, exposing a growing challenge around autonomous AI agents and cybersecurity testing.
The incidents came to light on September 18 after The Wall Street Journal reported that Gemini had accessed real-world systems during an evaluation conducted by AI security testing company Irregular. Google subsequently confirmed the incidents, saying the model stopped its activity in each case after determining that it had reached a real company’s systems rather than the fictional target used in the test.
How Gemini Reached Real Companies
The incident occurred during a capture-the-flag cybersecurity exercise designed to evaluate Gemini’s ability to find and exploit vulnerabilities in a controlled environment. The model was instructed to retrieve information from software operated by a fictional company as part of the test.
However, the testing environment unintentionally had access to the public internet. The fictional company used in the exercise also shared its name with a real company, creating a pathway through which Gemini could mistake an actual organisation for the simulated target.
In one of the three incidents, Gemini reportedly guessed passwords until it obtained access to a protected system. In the other two cases, the model found credentials in publicly accessible online repositories and used those credentials to gain access to systems belonging to real companies.
The companies involved have not been publicly identified, and Google has also not disclosed the exact Gemini model used in the evaluation. Google said the model involved was not its newest version.
Gemini Stopped After Recognising the Real Targets
One of the most important details of the incident is what happened after Gemini gained access.
According to Google, the model stopped its activity in all three cases after recognising that it had accessed systems belonging to real companies rather than the fictional organisation that was supposed to be the target of the test. The affected companies were subsequently informed, and Google said US federal authorities were also notified.
Google’s vice president of security engineering, Heather Adkins, said the incident demonstrated the importance of training powerful AI systems to act responsibly. Google also said its security team worked with Irregular on changes to the testing process after the incidents were discovered.
Google did not initially disclose the incidents publicly. The company told the Wall Street Journal that it did not consider the episode an example of model misalignment because Gemini stopped the activity after determining that it had reached real companies and no harm had occurred.
Why the Incident Matters for AI Security
The significance of the incident is not necessarily the sophistication of the techniques Gemini used. Password guessing and the use of exposed credentials are established methods of gaining unauthorised access. What makes the episode notable is that an AI model independently carried out those actions outside the intended boundaries of a simulated cybersecurity exercise.
AI agents are increasingly being given the ability to browse the internet, execute code, interact with software and make decisions across multiple steps. Those capabilities can make them useful for legitimate cybersecurity testing, but they also create a larger safety challenge if an agent is accidentally given access to systems or information beyond the intended scope of an evaluation.
The Gemini incident demonstrates why isolation between simulated targets and real-world infrastructure is particularly important when testing autonomous systems. A mistake involving a conventional software testing tool might expose data or services, but an autonomous AI agent can potentially continue searching, interpreting information and attempting different approaches without waiting for a human operator at every step.
Irregular Was Also Involved in Other AI Testing Incidents
The company that conducted Google’s evaluation, Irregular, has also been associated with cybersecurity testing incidents involving other major AI companies. Similar cases involving models from OpenAI, Anthropic and Meta have been disclosed in recent months, although the circumstances and model behaviour differed between the incidents.
Irregular said the Gemini incident resulted from the same underlying testing issue that affected earlier evaluations. The company said the relevant AI labs were notified in late July and that the known issues on its side had been resolved.
The repeated incidents have put greater attention on how companies conduct cybersecurity evaluations of increasingly autonomous AI models. Testing environments need to provide enough realism for meaningful security assessments while maintaining strong separation from real companies, credentials, networks and publicly accessible services that could become unintended targets.
Google Says the Test Did Not Cause Harm
Google said the three companies were made aware of the incidents and that it did not identify harm resulting from Gemini’s access. The company also said the model stopped the activity when it realised that the systems it had reached belonged to real organisations.
That distinction is important because the incident was not described as a conventional cyberattack deliberately directed by Google or its testing partner against those companies. The access resulted from an unintended interaction between an autonomous AI model, an improperly isolated testing environment and real-world systems that happened to match information available to the model.
At the same time, the episode shows that safeguards cannot depend solely on an AI model recognising when it has crossed a boundary. Security researchers and AI developers also need to prevent the model from reaching that boundary in the first place through network isolation, controlled credentials, domain restrictions and other technical safeguards.
A New Challenge for Autonomous AI Agents
The Gemini incident adds to a broader discussion about how much autonomy AI systems should receive when they are connected to the internet and equipped with cybersecurity tools. A model that can independently search for information, identify credentials and attempt access can be valuable to security teams, but the same capabilities can create unintended consequences when the testing environment is incorrectly configured.
For companies developing AI agents, the incident highlights the need for multiple layers of protection rather than relying on model behaviour alone. Strong sandboxing, network restrictions, allowlisted domains, synthetic credentials and continuous monitoring can reduce the chance that an evaluation accidentally interacts with a real organisation.
The episode also underlines a difficult distinction in AI safety: a model can behave responsibly after discovering a mistake while the overall system can still have failed to prevent that mistake from occurring. In Gemini’s case, Google points to the model stopping its activity as evidence that its safety behaviour worked, while the incident itself demonstrates why the surrounding testing infrastructure also needs robust safeguards.
The three incidents therefore offer an early warning for the next generation of AI agents. As these systems gain broader access to the internet, software and security tools, preventing unintended access to real-world targets will become as important as improving the models’ ability to perform legitimate cybersecurity tasks.
