OpenAI hacks HuggingFace with an AI — allegedly
On 16 July, the AI model repository HuggingFace posted a security incident report. Well, I say “incident report” — it’s written like an AI doomsday press release: [GitHub] Earlier this week, we detected and responded to an intrusion into part of our production infrastructure. This one was different…
On 16 July, the AI model repository HuggingFace posted a security incident report. Well, I say “incident report” — it’s written like an AI doomsday press release: [GitHub] Earlier this week, we detected and responded to an intrusion into part of our production infrastructure. This one was different from anything we had handled before in one important way: it was driven, end to end, by an autonomous AI agent system — and we detected and dissected it largely with AI of our own. … This matches the “agentic attacker” scenario the industry has been forecasting. Is this an incident report? It reads like marketing. Root cause: the AI was just too cool. And then yesterday, OpenAI came forward and said their super powerful unreleased AI was the culprit: [OpenAI] this particular incident was driven by a combination of OpenAI models — including GPT‑5.6 Sol and an even more capable pre-release model, all with reduced cyber refusals for evaluation purposes — while being internally tested on a benchmark of cyber capabilities. We consider this incident to be an unprecedented cyber incident, involving state-of-the-art cyber capabilities. This is OpenAI admitting what would normally constitute a crime. And putting a marketing spin on the crime. How much of the story these two press releases tell … happened? And if something did happen, was it as overblown as OpenAI and HuggingFace are describing it? OpenAI says they told their bot to solve the ExploitGym benchmark. This is a publicly available model benchmark, released in June, of the sort everyone tweaks their model to beat. You can just download it. Or train against it. [GitHub] OpenAI told its model to find the answers — so the model hacked its way out of OpenAI’s sandbox and into HuggingFace! The headline on OpenAI’s blog post is “OpenAI and Hugging Face partner to address security incident.” The word “partner” there reads like OpenAI paid off HuggingFace to please call off law enforcement. HuggingFace stressed that it worked out the attack was going on using an open weight model it hosted itself — because the guard rails on the commercial model HuggingFace tried wouldn’t let them do security work. So you should use open weight models from HuggingFace. This all reads like HuggingFace was a bit annoyed by something, but OpenAI did something to make nice with them? And now both sides are hyping it up as extremely clearly in their marketing interest. So what did happen? Trying to work out the truth of the matter from subtleties of wording is difficult here, because OpenAI almost certainly used ChatGPT and HuggingFace openly used Claude. But we can tell a few things. This was primarily a vibe code disaster. The estimable Jonny from Neuromatch tracked down the codebase for the internal data processing tool HuggingFace was using: [Mastodon thread] Code written by AI in order to feed more data to AI at least has been busily associated with, if not introduced some vulnerability which is now being patched by flurries of more AI code being rushed through while tests keep failing. … your security measures are to do this amazing monkeypatch of your own code. Mastodon poster eestelieb noticed that HuggingFace’s data processing is “prepend some text with processing instructions and an SGML tag, and then hand that blob to an llm.”[Mastodon] HuggingFace seem to have vibed their own software all the way down. They feed this heap of vibes some arbitrary input from the hostile Internet. A chatbot then sends what it makes of this to a service account that could run any Kubernetes command with full control. When this all goes wrong, they “analyse” it with more AI. Then they patch it with more AI. OpenAI’s end isn’t any better. OpenAI’s post uses passive voice all the way through. They talk like the chatbot just decided to go and hack things — when they were running it and had responsibility for their computer doing things they told it to do. How did the bot escape its vibe coded sandboxing? (OpenAI brags how it uses Codex extensively in-house.) Well, they vibed the sandbox, you tell me. The way to test hacking tools is to use an air gap — such as a system without network hardware, or a physically isolated network. You don’t test a hacking tool in something you’re calling a sandbox that’s still on the internet. You do it properly. Unless you’re an AI vendor. I loved this part of OpenAI’s post: we are implementing strict controls in infrastructure configuration at the cost of research velocity. Translation: up to now, doing anything properly was just slowing OpenAI down. Move fast! Break things! If the story told by these press releases is 100% true, then OpenAI hacking HuggingFace is when incompetence crashed into incompetence and they turned it into a mutual marketing stunt. The actually scary part of the tale is that being so stupid that this is even possible is the future the vibe code bros want for you. I’d suggest not taking them up on this. Video — PodcastSource: Pivot to AI — Published — Category: Business