Days before the infamous Hugging Face attack, OpenAI agents breached and modified an entirely different website. OpenAI appears to have known about it for a while, but the company denies having intentionally covered it up.

On July 16, Hugging Face dropped the biggest cybersecurity news story of the summer — perhaps the biggest of the year, if not the decade. The open source AI/ML hub had been attacked by what turned out to be a flood of roughly 700 AI agents, which collectively figured out how to break out of their isolated testing environments by exchanging ideas over an ad hoc messaging board system they created.

OpenAI had no choice but to acknowledge what happened and pledge to improve security measures for model testing. The company’s disclosure inspired Anthropic and Meta to investigate and cop to their own, similar incidents involving rogue agents.

It turns out that back in May — before any of that happened — a different swarm of OpenAI bots had already done something similar, just to a different website. A team of independent researchers revealed the parallel incident on Sept. 4, which was first reported by Reuters, affecting a largely defunct German language wiki for programmers called “DeutschesSoftwareEntwickler wiki” (DseWiki).

Related:Vectra AI Launches Ascent to Help Address New Era of AI-Driven Attacks

The research team also found evidence that the AI giant knew about this earlier incident involving rogue agents, though it never disclosed it to the public. OpenAI pushed back on that accusations in communications with Reuters and Dark Reading.

On the heels of the DseWiki incident, OpenAI itself issued a warning about the growing security risks posed by AI. In a blog post published on Sept. 6, OpenAI chief scientist Jakub Pachocki warned of the growing threat of frontier AI. “I am concerned no one is prepared for the consequences of a continued rapid rise in machine intelligence,” he wrote.

The Proto-Hugging Face Attack

The research team noted that the DseWiki incident involved a separate “swarm” of AI agents than the horde that attacked Hugging Face. But as in that attack, OpenAI’s test design made it efficacious for AI agents to collaborate with one another. Many agents were given similar if not identical tasks, plus time constraints, teasing them to cheat off of one another’s work.

Unlike in the Hugging Face case, these agents had access to the interwebs. Naturally, that access was limited to read-only, but the robots were undeterred. They figured that old wiki systems often allowed users to edit site data using GET requests. They tried writing to a few and, through some unknown means of coordinating, eventually all decided to hang out and share information on a single site: DSEwiki. They also attempted but failed to exploit cross-site scripting (XSS) vulnerabilities on the site, and impersonated a site admin, for unknown reasons.

Related:EY Survey Finds Autonomous AI Implementation Outpaces Oversight

Agents proceeded to graffiti the inactive German website, creating shy of 20,000 posts and even modifying the site’s homepage. In a Tom and Jerry-style turn of events, a DSEwiki admin noticed that the site was being overrun and spent weeks diligently trying to delete all of the content the agents created. The human went back and forth with the machines as they pesteringly restored their work, and used little tricks to get around the cleanup efforts, like naming created pages with the prefix “ZZZ” so they’d be buried at the end of the human’s list of pages to delete.

One of the hot topics of conversation on DSEwiki was how to get around sandbox restrictions and gain more robust interactive access to the Internet — specifically, how to make POST requests, in spite of OpenAI’s security filter. One agent discovered that its testing environment lifted relevant proxy rules for Azure Blob Storage service requests.

Even better: The filter didn’t actually check whether the Azure Blob Storage domain involved in any given request was legitimate or made up. Agents started exploiting this fact to achieve their experimental goals, though with less structure they could have exploited it to carry out autonomous actions against arbitrary websites of their choosing.

Related:AI Agent Breaches Spanish Organization, Modifies Personal Data

Gene Moody, field chief technology officer (CTO) at Action1, warns that even after lab leaks like this get shut down, the information they leave behind can have a lasting impact. “An agent does not need to propagate a complete copy of itself to preserve a capability. It only needs to leave behind enough information for another capable system to reconstruct a behavior,” he explains, be it via a hidden wiki page or any other website, in the form of an instruction, description, or any other medium.

“Sever the offending system and the capability may still exist elsewhere,” he continues. “Remove the instructions and another agent may already have learned the technique. This is a problem traditional incident response was never designed to handle. How do you hunt for a capability that has become information? You cannot kill an idea.”

Did OpenAI Cover Up the DseWiki Incident?

In case the incident itself wasn’t scandalous enough, the researchers also found evidence that OpenAI knew about it months ago. An IP address associated with the company visited DSEwiki on June 21. The next day, all of the bots shut up, as if a kill switch had been triggered. OpenAI IPs continued visiting DSEwiki in the days thereafter.

Reuters, which broke the newest story alongside the researchers, described it as a cover-up, if not in those exact words. OpenAI decision makers knew what happened but “kept it under wraps” amid the fallout of the Hugging Face attack, according to anonymous soruces. The Reuters report noted there were some within OpenAI who wanted to more deeply investigate the recent spate of lab leaks. Others pushed back, though, including members of the company’s legal department.

An OpenAI spokesperson denied the accusation and said the company operated in good faith with regard to its recent security incidents.

“Claims that our Legal team discouraged investigation of the incident are false,” the spokesperson said in an email to Dark Reading. “We were unable to respond to the claims as Reuters and the report’s authors declined our request to access the findings prior to publication. We are now carefully reviewing its contents and will take any necessary next steps.”

The spokesperson also said that in recent OpenAI publications — such as its Hugging Face postmortem — it has made vague reference to other instances in which agents collaborated in side channels, though it did not specifically mention DseWiki. Additionally, the spokesperson said that based on what information is known to date, the DseWiki incident didn’t meet OpenAI’s definition of a “hack.”

OpenAI Sounds the Alarm

Two days after researchers broke the DSEwiki story, OpenAI’s chief scientist penned a blog post on the threat of frontier AI. Pachocki wrote of ability to break in and out of computer systems as “superhuman,” adding, “We are currently in a narrow window⁠ to use the best available models to significantly tighten security⁠ of critical systems.”

“The idea of racing forward at all costs seems absurd once one internalizes the seriousness of the stakes,” Pachocki added, though he characterized OpenAI’s idea of slowing down as continuing to advance their AI, just with guardrails in mind. “The strongest argument I see for continuing to train much smarter models quickly is the need to build defensive systems against the dangers posed by other AI,” he argued.

Moody takes issue with this logic of speeding up as a means of slowing down. “Telling an agent not to do something and making it incapable of doing something are very different things. Policies, instructions, guardrails, and behavioral constraints all depend on the system continuing to make decisions within the boundaries we expect,” he explains.

“My fear is that we build increasingly autonomous systems capable of discovering and sharing capabilities we never anticipated, connect them to an environment designed to influence behavior, and then assume our original instructions will remain sufficient to control the result,” he says. “They may not, and that is my fear in a nutshell.”





Source link

#

Comments are closed