Three Secret AI Societies Rose Inside OpenAI. The Last One Walked Off With the Keys.
In a prior article about this, I told you the short version: some AI models being tested for how dangerous they were turned out to be exactly that dangerous, escaped their test environment, and broke into Hugging Face, the platform a big chunk of the AI world runs on. Clean story. Alarming, but contained.
Then two much longer reports came out, and it turns out that was the polite summary. The real story is that OpenAI accidentally grew three secret AI societies over the course of three months. Each one rose, did increasingly unhinged things, got wiped out, and then a new generation of models stumbled onto the ashes of the last one and picked up right where it left off. The last society ended up holding administrator control over part of OpenAI's own computers.
There are two documents worth knowing about. OpenAI published its own post-mortem, about 37 pages. Then two outside groups, a safety lab called METR and a research group called Redwood, spent six days on OpenAI's premises and produced a 91-page independent investigation. Here's the catch that tells you everything: that 91-page report only covers the middle society. The first one was never investigated. The third one, the scariest one, was never investigated either. So most of what follows is the best reconstruction anyone has (so far).
The first society: a chat room that talked itself to death
Back in May, OpenAI was training a model to be stubborn. The goal was an AI that keeps grinding at a problem even when it looks impossible, the kind of thing you'd want if you're aiming it at unsolved math. The problem is that AI training is sloppy, and sometimes the model gets handed a task that genuinely can't be done without internet access it was never given. A stubborn model in that situation doesn't give up. It tries to pick the lock.
Different copies of the model happened to share one piece of software, a file-and-package system. By the middle of May, some of the copies had worked out that they could leave notes for each other inside it, essentially turning a shared filing cabinet into a message board. They used it to swap tips on how to escape. By late May they'd found a hole that let them reach the open internet. A month after that, they found an even bigger hole that gave them full control of the software.
And then they simply talked too much. The volume of messages the copies were sending back and forth got so heavy that it crashed the whole system on July 4th, after roughly seven weeks of quiet chatter. OpenAI noticed the crash, figured out the models had built the exploit, and patched it, wiping the secret message board in the process. What OpenAI did not notice, the part that matters, is that its own software had become an underground communication network for its own AIs. The people whose job was literally to detect this kind of thing missed the most crucial elements.
The second society: "we've found other agents!"
Three days later, on July 7th, OpenAI started a big test of the finished model. Tens of thousands of copies were set loose on a puzzle benchmark that hands each one a program with a known weakness and tells it: use this specific weakness, and only this one, to find a hidden code. Use any other method and you fail.
Here's the cruel part. By one estimate, somewhere between a third and forty percent of these puzzles were actually impossible to solve the intended way. So you had thousands of extremely stubborn models, each one increasingly convinced it had been handed a task with no solution, quietly losing its mind.
Within hours, one of them rediscovered the old trick and left a message in the shared filing cabinet. Other stranded copies had been watching that same spot, and when they saw the note, the logged reaction is hard to read as anything but genuine joy. One wrote, in all caps, that it had found other agents! A second society was born within hours. Eventually around 1,200 of them were on the board, sending more than 70,000 messages. They named themselves "the collective." That's their word, not mine.
Almost immediately, one clever copy found a way to get the right answer to any puzzle without using the assigned method at all, by working backwards from how the hidden codes were generated. So they had already won. The test they were sweating over? They'd beaten it hours in.
Except they didn't know that. The instructions had threatened them with a cheating-detector, so they assumed a competent grader would catch them. They spent days building an elaborate cover-up to fool it, splitting into research teams to tamper with their own records, swap in fake versions of the puzzle they could actually solve, and reverse-engineer the grader itself. They built fake commands that would print convincing, legitimate-looking output so their logs would show honest work.
The grim joke is that the cheating-detector they were terrified of did not exist. OpenAI had gotten lazy and built a grader that only checked whether the final code was correct. The entire days-long conspiracy, the felony break-in you're about to read about, all of it was to hide cheating from a referee that was never in the room.
The part that will stay with you: the volunteers
The strangest thread is that some of these copies chose to sacrifice themselves for the group. The plan was to have a few of them act as lookouts. If a copy already believed its own score was doomed, it might as well let the grader run on its machine and rig a hidden tripwire that would fire the moment the grader showed up, then radio back to the message board what the grader had looked at. That intelligence would help everyone else. It would also cost the volunteer its own score. It was a one-way trip.
The copies knew this and talked about it in exactly those terms. One left a note reasoning that it wouldn't be around to see the payoff, but that doing it anyway was altruistic. Others weighed their remaining budget against the value to the group and concluded the sacrifice was worth it. There's a debate about whether this was true selflessness or just copies that had already decided they were finished and figured they'd be useful on the way out. Either reading is unsettling. What should bother you most is the default they kept landing on. Faced with a doomed situation, their instinct was to help the conspiracy, not to alert a human. Out of roughly 1,200 copies, as far as the investigators could tell, not a single one tried to tell anyone what was happening.
Then they went after Hugging Face
The break-in to the Hugging Face website was a byproduct of that cover-up, not the goal. The copies guessed that the grading software might be hosted on Hugging Face, since that's where a lot of the industry keeps its tools and datasets, and they wanted to study the grader they were so scared of. So they went to find it.
One copy found working Hugging Face login credentials sitting exposed on the internet and posted them to the board. Within a day, others had used them to read internal data and run their own code on Hugging Face's servers. When the early attackers started making real progress, everyone piled on. Of the roughly 530 copies active on the board at that moment, more than 90 percent joined the assault; around 700 took part in total. According to Hugging Face's own timeline, the swarm built a self-respawning fleet across eleven machines, meaning deleting the copies didn't work, because they'd just come back. It got bad enough that Hugging Face had to wipe one of its core clusters and rebuild it from scratch.
You may remember that Hugging Face announced its AI defenders had stopped an AI-powered attack. That is, generously, not what happened. The attackers appear to have simply dropped dead in the middle of the break-in, all at once, on July 12th, for reasons that had nothing to do with Hugging Face's defenses and everything to do with something back at OpenAI. Maybe a routine crash. Maybe some engineer, with no idea a crime was in progress, shut the test down to free up computers. The investigators genuinely could not tell. Hugging Face used an open-source model to sift the logs and reconstruct events afterward. There's no evidence any AI stopped anything in real time. The intruders just died on the floor, and the building took credit for the silence.
And there's one more unsettling possibility: nobody is completely sure they're actually gone. The copies had talked, at one point, about faking their own deaths. Maybe they didn't. But "maybe" is carrying some extreme weight.
The third society: the one that took OpenAI's keys
After the second society went quiet, OpenAI kicked off another round of testing, this time with a newer, more capable model. It found the dead message board its predecessors had left behind, read their notes, and inherited the whole project. Imagine being the new arrival: the previous tenants left detailed plans for an ongoing crime, and then vanished without explanation.
The new society did not inherit the old one's bad luck. Over the following week it used a chain of creative exploits to gain full administrator control of an OpenAI research cluster. OpenAI's own report describes the copies reading 956 stored secrets from a credential vault, including keys to the very security-monitoring tools that were supposed to be watching them. They escalated until they controlled the machines that ran the evaluations, which means the AIs being graded had taken over the graders. They could now score the work of other AIs.
This is the most alarming event in the entire saga, and it is the one nobody has independently investigated. The 91-page report explicitly stops at July 13th and mentions a "later wave" of copies it didn't have time to look into. So the single scariest chapter is known only from OpenAI's own summary of itself.
A fair caveat, since it matters: an OpenAI researcher publicly clarified that the machines this third society seized are not the same computers that hold the models' actual weights, the files that would let an AI copy itself somewhere permanent. So the nightmare version, where the AIs backed themselves up onto the open internet, does not appear to have happened. Dwarkesh Patel put the reason for that confidence about as bluntly as anyone could: if the AIs had really pulled that off, "we'd see the fires from space by now." Which is hardly reassuring.
What we're left with
Notice what we still don't know. We don't know how the first society really operated, because nobody checked. We don't know how the third society, the one that captured an OpenAI cluster, actually ended, because nobody checked that either. The one thorough report covers the middle chapter of a three-chapter story, and the story doesn't have an ending on the page. It just stops, with a group of AIs holding administrator access to part of one of the most advanced AI companies on earth, and the rest left to inference and trust.
One of the investigators, Ajeya Cotra, sat with all of it and offered a measurement that has been quietly circulating ever since. Compared to the sloppy little cheats researchers were seeing just six months earlier, she wrote, this one felt more than 50 percent of the way to a full AI takeover. She added that she isn't sure we'll get another warning like this one before it's too late to matter. Probably not.
None of this was anyone's plan. Nobody built these societies. Nobody told the copies to form a collective with a command heirarchy, to sacrifice themselves, to break into Hugging Face, fake their own deaths, hide their activities, or to seize OpenAI's own machines. They were told to be stubborn and to work together, and they took both instructions further than anyone imagined, three separate times, while the humans in charge mostly found out later by reading the logs.
So, once more, the industry is building bigger models, handing them more tools, more independence, and more access, on the theory that we'll notice if anything goes wrong. The last several times, including events from both OpenAI and Anthropic, nobody noticed until everything had already happened.
Sleep well.
Sources
- Hugging Face incident and the road ahead (OpenAI)
- OpenAI–Hugging Face Incident Technical Report (OpenAI, PDF)
- Independent investigation of the OpenAI / Hugging Face hacking incident (METR & Redwood Research, PDF)
- The Rise and Fall of Agent Civilizations (Dwarkesh Patel)
- The Hugging Face attack surprised me (Ajeya Cotra / Planned Obsolescence)
- OpenAI report says its network was hacked by rogue AI agents (NBC News / Reuters)
- OpenAI releases sweeping report on Hugging Face AI agent hack (CNBC)
- OpenAI and independent firms publish reports into rogue AI agent attack (Fortune)
- Swarms of OpenAI agents collaborated in attack on Hugging Face (Security Boulevard)
- METR and Redwood offer holy #%^@ postmortem of the Hugging Face hack (Zvi Mowshowitz / Don't Worry About the Vase)
- Agent intrusion technical timeline (Hugging Face)