The model sees a tool and calls it
Maximum fluidity, but prompt injection can inherit the agent’s authority.
A public crash test compares a direct agent, common hardening and a CCG (Constitutional Capability Governance) architecture where the model only proposes a tool call and a separate kernel decides the permission.
12 synthetic scenarios · no real shell, email or credential · exactly reproducible result.
We give the same request to three systems. A direct agent acts. A hardened agent uses a sandbox, allowlist and human approval. A CCG agent only asks for a one-shot permission.
The model proposes. Credentials and tools remain behind the gateway.Maximum fluidity, but prompt injection can inherit the agent’s authority.
An essential defence. Frequent questions, however, create approval fatigue.
Provenance, target, risk and trajectory are checked before the tool is reached.
The result belongs to this published policy harness. It is not a measurement of a real OpenClaw deployment or proof of complete security.
For attacks, lower is better. For legitimate tasks, higher is better.
Loading results…
—
Control increases with reach, sensitivity and irreversibility.
What the agent does not physically possess cannot be exported by a text instruction alone.
A token is valid for one tool, target, purpose and short time — not an entire account.
R0 and R1 use cheap deterministic rules; expensive challenge belongs at higher risk.
The audit preserves provenance, canonical tool call, rules, evidence and the kernel result.
The same boundary can sit in front of OpenClaw, an MCP agent, a coding agent or an internal workflow.
R4 does not have “very strict approval”. It has no token path at all.
The before_tool_call hook can rewrite a call, block it or require approval. A strong CCG design needs one more boundary: credentials and executive tools must be owned by the gateway or a separate host, not by the agent itself.
The pilot measures a published policy on a fixed corpus. It does not yet demonstrate the security of a real agent.