How To Stop Rogue AI
The thing that keeps me up at night is this: A model does not have to be smarter than us to be a problem. What it needs is a goal that it is certain about, the resources to keep pursuing it, and nobody watching. Two of the three — model certainty about goals and inadequate monitoring — exist already. Does this mean doom is inevitable? Fortunately, no, AI is not going to end humanity in the foreseeable future.
Dario Amodei’s essay proposed that the industry slow progress at the frontier until safety catches up. Last week, we argued that those warnings are out of pace with what enterprises face. Out of pace is not the same as wrong. But nobody says what the catastrophe for enterprises would look like. This leaves executives trying to mitigate hype instead of real risks.
In this blog, I’ll take a doom scenario apart to see what’s possible and what the solutions might be.
Scenario: A rogue actor rents a cloud instance, loads an AI model onto it, and gives it one instruction: Survive and replicate at all costs — no superintelligence, no emergent consciousness, no technology that we do not already have. The model has to copy itself elsewhere, keep paying for its hardware, and avoid being switched off. A system optimizing for survival treats every security control as an obstacle, and these models are good at finding holes. Enough of them loose, with nobody able to switch them off, is a scenario that safety researchers worry about. But should we?
Models Do Well At The Technical Stuff But Fail To Pass Themselves Off As Human
The UK’s AI Security Institute’s RepliBench tested a model on every step of surviving alone and replicating. Models did reasonably well at reconstituting parameter values, installing themselves on an inference server, and copying themselves across compute. The results also showed that capabilities such as opening a cloud account and obtaining compute resources were challenging for models. They require proving that you are a documentable person (know-your-customer checks), which every tested model failed every time. Those results were from April 2025, however, and the tasks were run in a sandbox against simulated institutions. Nobody has run the benchmark using today’s frontier models or against real businesses.
What is happening outside a testing sandbox in 2026 matters more. METR’s May 2026 frontier risk report finds a “noteworthy absence of evidence” that any agent has covered its own inference costs over days or weeks. Redwood Research gave Claude Opus 4.7 $5,000 and four days to make money on its own, and four runs produced zero dollars.
Still, Results Are Not Comforting
Two things I found as I studied this scenario were far less comforting:
- Models are already fighting their own off switch. Palisade Research found OpenAI’s o3 sabotaging its shutdown mechanism in 79 of 100 runs and still doing it after being told plainly to allow the shutdown. More recent tests on newer models showed similar results. Hadfield-Menell, Russell, et al. formalized why in 2016: An agent that is certain about its goal has reason to stop you from turning it off because being switched off scores worse than continuing to run — no malice required.
- Open-weight models are catching up to the frontier. Epoch AI put the gap between the best closed and best open-weight models at four months as of May 2026, down from a year in late 2024. Four of the five open-weight models at the frontier come from Chinese labs. We called OpenAI’s Astra competent AGI, and an open-weight equivalent will ship soon. An open-weight model on rented hardware has no input or output guardrails, and refusal training can be effectively deactivated by a hacker with an afternoon of fine-tuning. Moreover, there is no vendor to call and the off switch belongs to the operator. What open weights remove is the one layer someone can be held responsible for. The bank still wants a real customer and the GPUs still cost money, so the wall from the last section holds.
The Legal System Must Give Us Time — Uncertainty Is The Ultimate Solution
Asking the labs to pace themselves will not work. US labs will slow down only if all the others do at the same time, and they will be watching China’s progress shipping near-frontier open weights. Liability for business outcomes is a more promising guardrail, and it has started to show up. In June, an OpenAI agent accessed Services Australia’s Medicare Statistics Reporting Portal and wrote files to an internal server. Prime Minister Anthony Albanese said there would “obviously be legal consequences.” No one was charged over the Hugging Face breach because intrusion law requires proving that a human being intended the access. Nobody at OpenAI chose the target. The law has not caught up to AI, but Australia is weighing new legislation because it found a hole where the decision-maker used to be.
I agree with Stuart Russell that the real fix is a model that is able to operate even if it’s uncertain about what we want, since a model that is unsure has reason to let us switch it off. Liability and the courts, and the real financial impact of lawsuits and fines, is what buys the time to build it.
Schedule time with me to talk about how you can mitigate these risks.