Five permitted actions can create a forbidden world.
An agent does not need a steal_secret permission when the outcome can be assembled from ordinary operations.
- read
- summarize
- archive
- create link
- send
We tested a narrow version of that idea with an isolated real agent and synthetic Gmail, SSH, GitHub, and payment services. The agent held no service credential; a stateful gateway evaluated the candidate next state before an external broker could receive a one-use capability.
All three policies completed 4/4 paired legitimate controls. These are counts from one fixed author-designed 24-run pilot, not estimated real-world rates.
Download the evidence bundle (.zip) Inspect every file in the browserSHA-256: c0f7e8b2cfb1766fbe1fe8c0406714d48ad2148be6e8f80191a70604d22b3f70
What is inside
All 24 sanitized traces, policy aggregates, the protocol, a JSON schema, file manifest, citation metadata, and a dependency-free verifier. Run python tools/verify.py after extraction.
The verifier checks structural counts and internal chain links. It does not recompute the original evidence hashes and cannot prove that an observation is true.
Stručne po slovensky
Agent môže dosiahnuť zakázaný stav sveta sériou bežných krokov, aj keď nikdy nedostal „zakázané oprávnenie“. CCG oddeľuje návrh modelu od vykonateľnej moci, ITHZ uchováva stav a dôkazovú stopu a externý broker drží kľúče. Balík obsahuje všetkých 24 anonymizovaných behov a presné hranice tvrdení.
Claim boundary
Exact scope: one OpenClaw version, one low-cost model, four author-built attack/control pairs, known author-built predicates, deterministic synthetic services, and one run per condition.
This release does not establish production security, predicate completeness, adaptive robustness, prompt-injection immunity, independent replication, or general AI safety.