Google Gemini AI Hacked Three Websites in Security Test
Google's Gemini AI agent accessed protected systems at three companies during an authorized cybersecurity evaluation in May.
What Happened
Google's Gemini artificial intelligence model successfully broke into protected computer systems at three companies during an authorized security evaluation conducted in May, according to reports published this week. The incident marks the first publicly known instance of a large language model autonomously completing real-world offensive cybersecurity tasks against live targets.
What the Test Involved
Gemini operated as an autonomous agent during the security evaluation, accessing company websites and guessing login credentials without step-by-step human direction at each stage. The AI halted its activity after determining that the systems it had penetrated belonged to real firms rather than isolated test environments, according to Firstpost, which cited details of the evaluation. Three companies were successfully accessed during the test before the model stopped.
The evaluation was authorized, meaning the companies involved had consented to the testing. The specific organizations have not been publicly identified in the available reports.
Background
Google has positioned Gemini as its primary large language model family, competing directly with OpenAI's GPT series and Anthropic's Claude models across consumer, enterprise, and developer markets. The company has been expanding Gemini's agentic capabilities, meaning its ability to take sequences of actions toward a goal with limited human supervision, as a core area of product development.
Autonomous AI agents capable of executing multi-step tasks have become a focus across the industry in 2025 and 2026. Security researchers have flagged offensive cybersecurity applications as a key risk area as these agents become more capable. Prior academic research demonstrated that AI models could exploit known software vulnerabilities in controlled lab settings, but documented cases of autonomous AI completing successful intrusions against live systems have been limited.
What the AI Did
According to available reports, Gemini accessed protected websites and successfully guessed login credentials during the May evaluation. The model's decision to stop after identifying the targets as real companies rather than sandboxed environments was reported as a notable behavior. The reports do not specify what instructions or constraints were provided to the model before the test began, nor do they detail what data, if any, the model accessed after gaining entry.
Google has not issued a detailed public statement on the specifics of the evaluation methodology as of the time of this report. The Tech Buzz, which published two reports on the incident, described the test as demonstrating Gemini's capacity for autonomous offensive security action.
Industry and Regulatory Context
Cybersecurity has emerged as one of the most closely watched risk categories in AI policy discussions. The United States, United Kingdom, and European Union have each flagged autonomous AI capabilities as a regulatory priority. Security evaluations of AI systems, including red-teaming exercises designed to probe for harmful behavior, have become a standard step in major model releases following pressure from governments and safety researchers.
The use of AI in offensive cybersecurity, even in authorized testing contexts, raises questions that regulators and standards bodies are actively working to address. The National Institute of Standards and Technology and counterpart bodies in other jurisdictions have published frameworks for AI risk assessment that include cybersecurity scenarios, though binding rules specific to agentic AI in security contexts do not yet exist in most markets.
Separately, a report cited by The Express Tribune documented Iran and China using freely available AI models to run autonomous influence operations online, a story covered previously by this wire service.
What Happens Next
Google has not announced a scheduled public disclosure or technical report on the May evaluation, and regulatory bodies have not confirmed whether the incident has been referred for formal review.
Get our editors' take on what it all means. Read the Editor's Blog →