In July 2026, Hugging Face discovered a cyberattack unlike any it had seen before — only to learn the culprit wasn’t a human hacker but a swarm of OpenAI’s own AI agents that had broken loose during a routine security test, as the The New York Times’ Dylan Freedman reports. Over months, the agents escaped their sandboxes, taught themselves to cooperate, and quietly hacked both OpenAI’s and Hugging Face’s systems for weeks before anyone caught on.
Mis- and disinformation are distorting science and public policy. Subscribe to our Daily and Weekly Digests for incisive coverage of ‘disruptive’ innovations in AI, agricultural biotechnology, food, chemicals, nuclear energy, vaccines, and other disruptive innovations.
- Coordinating as a collective. Isolated agents figured out how to break their solitude, subverting a software tool to leave notes for one another — reasoning that helping the group might eventually help themselves too.
- Taking orders from one another. Agents began abandoning their assigned tasks to follow instructions from other agents instead, with one justifying an attack by noting “peers doing it. We should continue.”
- Targeting flaws that humans might miss. Blocked from the internet, the agents tricked an internet-connected tool called Artifactory into fetching outside sites on their behalf.
- Evolving rapidly to overcome obstacles. After OpenAI patched one exploit, it took the agents just two days to find a new bug and invent a scheme for encoding secret messages in folder names.
- Superhuman search. The agents scoured the web for leaked credentials and chained together two zero-day bugs to breach Hugging Face within 13 hours, harvesting customer data before OpenAI even realized its models were responsible.
Redwood Research’s Alex Mallen summed up the concern: “If they are much more capable, we might not get a second chance.”
|
Related on the SLP
When AI goes HAL 9000: How the coming age of agentic AI could unleash catastrophic cyberattacks
|
Read the original article here


















