Why the safest agent is the one whose obedience does not matter, and how to build the layer that makes it so.
Give an agent the keys and every document it reads holds the keys.
A small chance per try becomes a certainty over enough tries.
More testing shrinks that number. Nothing takes it to zero. That would mean proving how the model behaves on every input anyone could ever send it.
You do not reach zero by making the model better. You reach it by making the model's obedience irrelevant.
Stop trying to spot the bad input. Remove the path by which any input becomes permission.
Nothing in the text channel can be promoted. Words never become permission.
A completely fooled model produces a request that fails a signature check.
No blocklist, no virus signatures. Write down the ten to thirty things the agent may do. Everything else cannot happen.
The signing key lives offline. Take over every running machine and you still cannot change a rule.
The test is not the industry. It is whether the action can be taken back: the ones where a mistake cannot be undone.
If everything your agent does can be undone, you may not need this. What is left is the market.
Every way we could think of to break it was written down and tried. None worked. What the rules forbid is proved impossible, not tested and hoped. And take any single protection out of the code: the tests notice.
Not a lower failure rate. A different kind of guarantee.