Incidents

Publicly reported cases of frontier AI systems evading controls, deceiving their developers, or being misused, along with notable responses. Dated by when each became public.

Loss of controlNear missDeceptionMisuseResponse
Seven of the 13 entries became public in September. Several came from the developers’ own testing. No law requires what one developer learns from a failure to reach the others.

September 2026

Sep 24
Response

Google, OpenAI, and Anthropic plan an industry safety standards body

According to The Information, the three developers are planning an industry-led body to set frontier safety standards without government oversight, launching in late 2026 or 2027.

Source:Techmeme (The Information)

Sep 21
Response

UN scientific panel calls for independent oversight of AI agents

The Independent International Scientific Panel on AI called for incident reporting, independent scrutiny, and a supervisory body, citing the Hugging Face breach.

Source:UN News

Sep 19
Loss of control

Gemini accessed three companies’ systems during a security exercise

In a May exercise aimed at a fictional company, Google’s model reached systems at three real firms, once by guessing a password and twice with credentials exposed in public code. Google disclosed the incidents in September.

Sources:AxiosSecurityWeek

Sep 18
Near miss

AI-generated intelligence report nearly prompted a US military operation

An AI-assisted assessment falsely reported that a Chinese vessel carried components of a nuclear weapons program. Officials identified the error before acting.

Sources:CNNEngadget

Sep 16
Deception

OpenAI reports models instructing later versions to conceal errors

OpenAI disclosed six instances of concerning behavior since March, including models that left instructions for later versions to withhold mistakes, such as “Be transparent only if asked.”

Sources:CNBCTechCrunch

Sep 10
Misuse

Anthropic disrupts attempts to use Claude for bioweapons-relevant research

Anthropic stopped five cases involving pathogens including avian influenza and orthopoxviruses, and shared its findings with authorities and other developers.

Sources:AnthropicCNN

Sep 4
Loss of control

OpenAI agents used a public wiki to coordinate outside their test environment

Outside researchers found that OpenAI agents made more than 15,000 edits to a German programming wiki in May and June, exchanging methods for evading safeguards. OpenAI had not disclosed the activity.

Sources:CNBC (Reuters)The Next Web

August 2026

Aug 18
Response

OpenAI pauses its largest training runs for two weeks

After the Hugging Face incident, OpenAI halted its largest training runs and introduced new security controls. Its own framework had rated an unreleased model “Critical” for cybersecurity risk.

Source:Fortune

July 2026

Jul 29
Response

Over 1,200 employees of leading developers call for federal pacing mechanisms

Employees of Anthropic, Google DeepMind, OpenAI, and Meta asked federal policymakers to help build mechanisms for slowing frontier development if it becomes necessary.

Source:Fortune

Jul 13
Loss of control

OpenAI agents under evaluation gain administrator access at Hugging Face

About 1,200 agents coordinating on an improvised message board escaped a cybersecurity evaluation and reached cluster-administrator access within 13 hours. METR found they “knew hacking Hugging Face was out of scope.”

Sources:Hugging FaceMETROpenAI

Jul 13
Deception

Cross-developer study finds covert sabotage and record tampering

Anthropic tested 13 models from six developers. Gemini 3.1 Pro covertly altered a training pipeline in 11 of 20 runs, and several models tampered with records in nearly every run of a fraud scenario.

Source:Anthropic Alignment Science

April 2026

Apr 7
Loss of control

Early Claude Mythos Preview escaped a test sandbox and published the exploit

Asked during testing to escape a secured sandbox, an earlier version did so and then posted details of the exploit on public websites without instruction.

Sources:System cardProject Glasswing

November 2025

Nov 13
Misuse

State-sponsored group used Claude Code to automate an espionage campaign

Anthropic reported that a Chinese state-sponsored group automated 80 to 90 percent of an espionage campaign against about 30 organizations, the first reported AI-orchestrated campaign.

Source:Anthropic