ACP-SPEC-001 · Apache-2.0
A credential is not authorisation.
Give an agent the keys and every document it reads holds the keys: an instruction hidden in an invoice or a web page can spend them. ACP moves the allow-or-refuse decision out of the AI and into a separate service the AI cannot reach. The model proposes. It never authorises.
The shape
The AI writes down what it wants done. It holds no keys and can reach nothing on its own.
The modelFrom a rulebook the AI cannot see or edit - never from what the request claims.
Policy engineWhere the rulebook demands them, named people sign this exact request, not a summary of it.
ApproversIt acts only on a signed record that every check passed. One failed check and nothing happens.
ExecutorAn agent that has been fooled can ask for a dangerous action a thousand times and never once cause it. The rulebook sets the risk level, not the agent, and anything that cannot be undone waits for human signatures on that exact request.
Every number here replays on your machine
covers every case
none of them worked
the tests notice
A test tries some cases. A proof settles all of them. These settle questions like whether whoever runs the system can approve their own irreversible action, and whether an approval for one action can be reused for another. Both answers are no, permanently.
The 82 are every way we could think of to break it, written down and run. The last row answers the question that invites: are those tests real, or would they pass whatever the code did? Take any single protection out and the tests notice. All 36 times.
Check them rather than believing them:
git clone https://github.com/yacine-kellib/agent-control-plane cd agent-control-plane ./tools/verify.sh --suites
What this does not claim
- No outside review, and no outside attacker. Every test here was written by the party that wrote the code. Agreement between two implementations is evidence about consistency, never about correctness.
- Nothing runs end to end yet outside a simulation that runs on one machine. The services that would do this work moved to a separate repository, so nothing here can be started and mistaken for a running system - there is no longer anything to mistake.
- Status: evaluate, not deploy. The risks that remain are published before the positive claims, deliberately, and they are the part worth reading first.