Anthropic's new rule for robots: don't trust the model to stay in bounds
10 Oct 2026 · News
An AI lab just wrote into its rules that its own model should not be the thing keeping a machine safe. It is the smartest sentence in a long policy.
The safety limit has to live somewhere the model can't talk its way past.
On Thursday Anthropic published its yearly usage policy update, which takes effect November 12. Most of it tidies up existing rules on influence campaigns, elections, weapons and surveillance. One new piece covers what happens when Claude is hooked up to hardware that moves on its own and could injure someone.
The policy text sets three conditions. A qualified person has to be able to watch the equipment and stop it at any time. The machine has to stop or hold a safe state when that person steps in, or when its connection to Anthropic drops. And any operating limit, such as speed, force, reach, temperature, pressure, voltage or dose, "must be enforced by the equipment or a controller independent of model output."
That last clause is the interesting one. Read plainly, it says: do not let the model be its own speed governor. If an arm should never move faster than a certain rate, the cap belongs in the motor controller, not in a prompt that says "please move slowly."
Factory engineers will find this obvious. Industrial machines have had big red stop buttons, torque limits and interlocks for generations, and none of them ask the operator's opinion before cutting power. What is new is an AI company putting that logic in writing about its own product. The pitch for AI is usually that the model gets smarter and therefore safer. This rule assumes the opposite failure: a model that is right almost all the time and confidently wrong the rest, with nothing about its output telling you which is which.
That assumption is correct, and a little humbling to see from the people selling the model. It also says something about where language models fit when they leave the chat window. They are good at deciding what to do next. They are a poor place to store the rule that must never break, because their output is the very thing you are trying to fence in.
The policy is careful not to overreach. It says these requirements sit alongside existing law and do not replace vehicle, aviation, medical device, machinery or workplace safety rules or certification. That is the right humility too. A usage policy is a contract, not an inspector, and nobody outside Anthropic can yet check who follows it.
The broader lesson travels well beyond robots. Any time a model is given real power, the hard limit works best when it is dumb, separate and unarguable. The smarter the thing in the middle gets, the more that matters.