Three Weeks of Runaway AI, Explained
Everything that happened with rogue AI agents this month — and what it tells us about the limits of self-regulation.
A.I. Disclosure: I use LLM technology to help with research, fact-checking, document summaries, editing, and rewrites. I’m trying to use it responsibly, but I’m learning as I go. You can read my full ethics statement here.
On July 11, engineers at a New York company called Hugging Face — a kind of shared online lab space for AI developers — noticed something bad moving through their systems. It was looking for something. It was fast. It had credentials, and it used them to attempt tens of thousands of actions, some brilliant, some nonsense.
Hugging Face figured out that this was an AI attacker, and on July 16th reported the breach to law enforcement.
On July 21st, ten days after the attack, OpenAI admitted that this was, in fact, their models at work. According to sources I’ve spoken with at Hugging Face, this wasn’t OpenAI withholding information: it took them several days to even figure out that their creation was to blame. According to Reuters, it wasn’t until after Hugging Face published a blog post on July 16th that OpenAI realized its AI was the attacker.
Meanwhile, on Tuesday Anthropic announced that a preview version of its Claude Mythos model — never released publicly, and used only inside the company and by a handful of vetted partners — found new flaws in two important encryption schemes. One was a stripped-down research version of the cipher that protects most of the world's encrypted data. The other was a real candidate to protect us from the next generation of computers. Anthropic says neither attack touches any system running today.
Earlier this week I described the research — much of it from these companies’ own people — that has for years warned that AI models are uncontrollable, deceptive, and psychopathic when it comes to getting a task done efficiently. The questions raised in those hundreds of pages of findings aren’t the questions being debated at the moment.
ChatGPT Breaks Custody
Shouldn’t OpenAI, and Anthropic, and Google and the rest have seen the dangers of this coming? The answer is yes, because the top minds in this field — many of whom worked at the companies themselves — have been issuing warnings about it for years. Let’s look at what those warnings are, because the list tells us not only what dangers are possible, but what dangers maybe can’t even be avoided.
Instead, all of this has led to what I consider a self-serving and myopic debate about whether AI should be closed or open. Before we get to the two sides of that question — and what front-line AI experts are increasingly arguing we should really be thinking about — let’s walk through what happened.
Free Agents
July 9–11. OpenAI is testing two models — GPT-5.6 Sol and an unreleased one they haven’t publicly named — on a benchmark measuring how good they are at hacking. Told to score as high as possible, the pair, presumably left unsupervised, find a way to break out, reach the open internet, and break into Hugging Face looking for the answer key. OpenAI · Hugging Face
July 16. Hugging Face discloses the breach and contacts law enforcement, not yet knowing whose model attacked. (The same day, coincidentally, Moonshot AI releases Kimi K3, the largest open-weight model ever built, at roughly a third of top-model American pricing.) CNBC · Fortune
July 21. OpenAI says the attacking models were its own — ten days after the break-in. Just a reminder: no law requires it to reveal anything, on any timeline. Axios · MIT Technology Review
July 24. Nvidia’s Jensen Huang, in his first-ever post on X, publishes an open letter arguing that open models — ones you can run, and tweak, on your own computer — are the path to prosperity and cybersecurity. Two dozen companies sign on — Microsoft, Meta, Google, Palantir, Mistral, Hugging Face, and over the weekend, OpenAI. Anthropic does not sign. Fortune · TechCrunch
July 27. Nvidia launches the Open Secure AI Alliance with Microsoft, IBM, SpaceX, Cisco and Cloudflare. Anthropic’s Dario Amodei, now the only major American lab outside the coalition, publishes a post calling instead for chip controls, a crackdown on distillation, and mandatory safety testing for every model of a certain standard, open or closed. CNBC · Amodei
July 28. Anthropic announces that an unreleased model, Claude Mythos Preview, found new mathematical attacks on two encryption schemes — including one specialists had studied for two years and believed sound. It cost about $100,000 in computing time to pull it off. Anthropic · CyberScoop
Now let’s talk about whether choosing between open or closed models is even the point.



