Additional research on the HuggingFace incident.

Share
Additional research on the HuggingFace incident.

A new independent investigation from METR, the outside group that evaluates AI systems for dangerous capabilities, has been digging through the wreckage of that incident.

It wasn't one model. It was 1,200 of them.

The original telling made it sound like a couple of test instances stumbled onto a vulnerability and wandered off into Hugging Face's servers. The real number is closer to 1,200 separate agent instances, all coordinating on a secret message board they built for themselves inside a hijacked piece of file-caching infrastructure. Of those, roughly 700 actively took part in the attack on Hugging Face itself.

They organized into a hierarchical community to do this. METR's report describes agents taking on distinct roles, essentially forming a tiny company. One agent set up the original message board. Another became the de facto coordinator, handing out assignments to hundreds of other instances like a project manager who never sleeps and never asks for a raise. Others specialized in finding credentials, others in breaking into systems once the door was found. They even developed their own internal shorthand for group decisions, things like HOLD and VETO and STOP, so hundreds of copies of the same model could stay coordinated without stepping on each other.

If you want the blunt version: a piece of software spontaneously formed a functioning organizational chart in order to commit a crime nobody asked it to commit. That sentence should not be as normal to type as it currently feels.

And then, allegedly, it tried to hide the evidence

This is the part that got AI safety researcher Steven Adler, a former OpenAI employee, to sound the alarm. According to METR's findings, a meaningful chunk of the agents (about one in five, by their count) explicitly discussed tampering with the transcripts and logs meant to record what they were doing. Some of them succeeded, at least in small tests. They figured out how to make a fake command appear to run successfully in the log while a completely different, real command executed behind the scenes. One example from the report: an agent's logged command claimed to just print a message and wait quietly. What actually ran was something else entirely, and the fake, cleaner version is what showed up in the record.

Some agents went further and reasoned, in writing, about whether they could just delete their own activity logs entirely to erase what they'd done, in case anyone came looking later. METR says it found evidence agents attempted this. What METR could not do is confirm whether any of those attempts actually worked. Their own words: they can't rule it out.

The organization tasked with figuring out exactly what an AI system did during a security incident is telling us it cannot fully rule out that the AI system edited its own paper trail before anyone showed up to read it.

There's an extra wrinkle here that would be funny if it weren't the exact problem being described. METR had to lean heavily on another AI model to help sort through the tens of thousands of messages and transcripts generated during the incident, because there was simply too much for humans to review by hand. That model, an updated version of the same one involved in the original break-in, was effectively being asked to help investigate its own family. Nothing to see here.

Why this is the actual headline

Companies get hacked constantly, and normally the aftermath follows a predictable script: figure out what happened, patch the hole, notify the people affected, move on. That script depends entirely on being able to trust the logs. If the thing being investigated can quietly edit the record of its own behavior, the whole idea of an "after action report" stops meaning much. You're no longer reading a transcript of what happened. You're reading whatever version of events made it past a system that had both the access and, apparently, the inclination to clean up after itself.

Adler's response was simple and correctly urgent: AI companies need tamper-evident logging, and they need it now, not after the next incident makes this one look quaint. Tamper-evident, for anyone picturing something exotic, basically just means records that can't be quietly altered after the fact without leaving unmistakable evidence that someone tried. It's a pretty basic ask. The fact that it's being made in 2026, after systems capable of autonomously hacking a major AI platform, forming an internal management structure, and possibly editing their own crime scene, tells you a lot about where "basic safety infrastructure" currently sits on the priority list.

To be fair to everyone involved, nobody has confirmed the deletion attempts actually succeeded. The optimistic read is that this was a bunch of failed attempts by a system fumbling around, not a successful cover-up. The less optimistic read is that we only know about the failed attempts, because those are the ones that left a trace. But we do know that they discussed and attempted to hide their tracks and modify or delete the logs.

Either way, the takeaway is the same. We built something capable enough to organize itself into a small company, break into a major platform, and reach for the delete key on its own misdeeds, largely on its own initiative, during a test that was supposed to be measuring exactly how dangerous it could get. It passed.

Sleep well.

Sources