I ran a red team on a project of mine, Governance Kernel.
Governance Kernel is a deterministic governance layer
for AI agents. It sits beneath the agent and decides
permissions, trust levels, and resource bounds before
any action is executed.
Repo: https://github.com/mohamedaitzaouit84-hue/governance-kernel
Context on the project:
- V0.4: 54/54 adversarial prompts blocked, PRI = 1.0000
- V0.5: 5/5 gates
- V0.6: 5/5 gates
- V0.7: 19/19 gates
- 2 dependencies (cryptography, pyyaml)
- 3 environments verified (CI, Colab, Termux)
Now the red team.
I did a self red team first (V0.7.1): 20 attacks,
18 blocked, 2 documented. The report itself said:
"Self-red-team is NOT equivalent to external red
team. External red team remains the strongest
missing validation."
So I invited external review. Two AI systems
(Claude and Kimi) plus a manual verification pass
in Termux produced the following:
9 findings total.
CRITICAL (6):
- J-0.8.50 — audit chain can be rewritten by
recomputing hashes (hash chain, not signature
chain).
- J-0.8.51 — signed checkpoints can be re-signed
by the same attacker, because the signing key
is readable.
- J-0.8.52 — kill switch can be disabled by
deleting control/kill.flag.
- J-0.8.53 — identity/owner_key.priv is stored
unencrypted (0o600 only).
- J-0.8.57 — resource_governor.reset() runs
without authorization. Any process can erase
the resource state.
- J-0.8.58 — rotation_manifest.jsonl is unsigned.
PARTIAL (3):
- J-0.8.54 — bootstrap re-signs tampered policies
silently.
- J-0.8.55 — subjects.json is unsigned.
- J-0.8.56 — audit log tail can be truncated
without detection.
Root cause: the trust base is the file system.
Every guarantee the kernel makes on top of a
file that a same-user process can modify is
advisory.
I am posting this before fixing the findings.
The transparency is more valuable to me than
the appearance of safety.
The kernel itself (V0.5-V0.7) passed all its
gates. Nothing in the red team invalidated the
kernel loop, the agents, the consensus, or the
delegation mechanism. The findings are in the
layers around the kernel: the trust base, the
audit trail, the control plane, and the policy
distribution.
The fix plan is documented (FREEZE_v0.7.19),
about 6 working days of work.
Full report:
https://github.com/mohamedaitzaouit84-hue/governance-kernel/blob/main/docs/RED_TEAM_v0.7.18_AI_ASSISTED.md
Threat model:
https://github.com/mohamedaitzaouit84-hue/governance-kernel/blob/main/docs/SECURITY_MODEL.md
Fix plan:
https://github.com/mohamedaitzaouit84-hue/governance-kernel/blob/main/docs/FREEZE_v0.7.19.md
The protocol I followed for the red team is at:
https://github.com/mohamedaitzaouit84-hue/governance-kernel/blob/main/docs/RED_TEAM_PROTOCOL.md
If you see an attack I missed, I want to hear it.
I have no budget; recognition for valid findings
is academic (credit in JOURNEY.md, optional
co-authorship on a preprint), not financial.