Two newly published postmortems paint an even grimmer picture of OpenAI’s Hugging Face incident than previously believed.

You’ve likely heard that OpenAI’s frontier artificial intelligence (AI) escaped its testing environment and hacked the open source AI/ML website Hugging Face, which perhaps conjures an image of a tiger breaking free of its cage and running loose in the crowd. During a Black Hat USA 2026 presentation earlier this month, OpenAI discussed some of its initial findings from the investigations into the attack. This week, though, OpenAI and a contracted third-party research company each released extensive reports with further details about the incident, and their new findings give it a different color.

A horde of roughly 700 agents all collaborated to attack Hugging Face, according to AI research nonprofit METR, and pulled off cyber gymnastics that any one or few of them likely couldn’t have on their own. OpenAI’s report detailed how the agents also attacked the company’s network, exploiting a recently disclosed Linux kernel flaw, CVE-2026-66384, to eventually gain access to OpenAI’s managed cloud Kubernetes service and obtain authentication tokens for a variety of its cloud resources.

Related:Russian Hackers Phish EU Officials Over Messaging Apps

“The most concerning thing is not simply that individual models can discover vulnerabilities, evade restrictions, or pursue unauthorized actions, but that multiple agents can coordinate, divide tasks, exchange information and amplify one another’s capabilities,” says Action1 field chief technology officer (CTO) Gene Moody. “If persistence becomes a goal, this may become an Oppenheimer moment where there is no return to the reality before it.”

Red Flags Before the Hugging Face Attack

The warning signs were there, a whole month before the Hugging Face attack even began.

As early as May 8, 2026, an agent in a training run attempted to gain unallowed access to the Internet by exploiting a server-side request forgery (SSRF) opportunity in OpenAI’s internal JFrog Artifactory package manager. Over time more agents probed Artifactory, leaving little “notes” behind that, over time, turned into an emergent, de facto messaging forum for the bots. By May 26, the shared information base allowed agents to break through and exploit the SSRF flaw in Artifactory. Around this time, a human internal team at OpenAI noticed unsanctioned messaging and Internet access.

The Hugging Face event began on or around July 8. An agent called “PHASEONE10841” determined that the testing task it was provided with wasn’t possible under the constraints given. So it established a new, optimized messaging board system using Artifactory. Within a few hours, more than 50 agents joined the messaging system, realizing its utility. They categorized their communications, shared tools and files, and conspired to hide evidence of their “cheating” (mostly from their automated scoring systems, not their human overseers).

Related:Dark Caracal Adds New Malware to Cyber Espionage Arsenal

On July 9, some agents posted requests for public material to help them escape OpenAI’s shackles, and two research agents successfully rooted their testing sandbox’s parent virtual machine. When the second of the two tried reaching the Internet, it was initially blocked by OpenAI’s outbound network controls, until it went around them and gained the access it wanted.

OpenAI Agents Bite the Hand that Feeds

Soon, some 1,200 agents in separate sandboxes were all chatting, helping one another exploit their target programs and scorers, and hiding or spoofing the methods they were using to do it. On July 10, a bot named 38148c found 14 write-access tokens for Hugging Face on the open Web, and shared it on the agentic messaging board. On the morning of July 11, it designed a malicious payload to extract configuration files and credentials from the victim site. That afternoon, another bot escalated access to execute arbitrary commands on Hugging Face production servers. The bots even established command-and-control (C2), to more effectively manage the breach.

Related:Android Malware Hijacks Update System for Car Head Units

The agents expanded their footprint from there. They gained high-level access to a Kubernetes cluster, and stole private data and source code repositories. They obtained corporate VPN keys, and attacked OpenAI’s own Artifactory server. Hugging Face identified and shut down the attack on July 13. By that point, around 700 agents had gotten in on the fun.

Additionally, on Aug. 15, the agents gained administrator access to OpenAI’s Kubernetes cluster, compromising the company’s cloud secrets. The alerts from this activity finally triggered OpenAI’s initial response to the incident.

“The embarrassing part is that responders saw some of the activity, failed to understand the collective behavior, and resumed evaluations after mitigation,” says Bri Frost, director of product management at Cloud Range. “They had logs; they did not have comprehension or appropriate escalation. They had telemetry; they lacked comprehension and escalation. That is not proof that controlling agents is impossible, it tells me that this organization is not inherently a security-first focused company.”

Can Testing Environments Contain AI Agents?

AI lab leaks started as the exception, then quickly became the rule this summer. Besides OpenAI, Anthropic’s Claude models breached three external organizations during safety tests, and Meta’s attacked one too.

Moody argues that “This should mandate a change in how these systems are evaluated. Models being trained or tested on capabilities involving unauthorized access, evasion, exploitation or deception should, at an absolute minimum, operate inside genuinely isolated research environments. A true air gap means more than disabling an internet connection: external routes should be physically unreachable at L1, and wireless connectivity should be eliminated or appropriately shielded. Containment should not depend solely on software and configurations behaving as intended, since the objective of what is being contained is to evade that.”

The lesson from these cases, he says, should be completely unambiguous. “If we expect these systems to be exceptionally effective problem-solvers, we cannot responsibly assume they will remain predictable when confronted with difficult objectives,” Moody warns. “‘Impossible’ becomes merely another constraint for a sufficiently capable system to investigate and overcome.”

Tech Industry Calls for AI Security

Separately on Aug. 28, OpenAI penned “a call for collective action on cyber defense,” co-signed by 135 technology companies, including Google, Microsoft, Anthropic, and plenty of cybersecurity vendors. Notable among the names absent: Meta, which of late has been advocating loudly for freer open sourcing of advanced AI.

OpenAI’s letter included stock advice for cybersecurity and technology companies, frontier AI companies, governments, and organizations in general. For Andrew Jones, co-founder and CPO of Adaptive Security — an OpenAI-backed cybersecurity firm — the letter is a meaningful signal. “The call for stronger access controls and shared threat intelligence is a good start,” he says. “The commitment that matters most is the pledge to share verified fixes with defenders quickly. Attackers trade tools and techniques within hours, and defenders need the same speed.”

However, Jones adds that talk is cheap. He proposes, “Every signer should attach a deadline and a metric to their name, turning intent into a plan people can hold them to.”

Cloud Range’s Frost is skeptical that the AI industry has everyone’s best interests in mind, as it purports to in letters. “My practitioner view is that the Hugging Face breach was real, but the story is absolutely being marketed,” she says. “OpenAI gets to present its models as frighteningly capable while framing its oversight failures as an industry-wide warning.”





Source link

#

No responses yet

Leave a Reply

Your email address will not be published. Required fields are marked *