AI Governance
Cowboy constrains what an agent can do through several independent mechanisms. They overlap deliberately: a command rejected by one layer is not relied upon to be caught by another.
| Layer | Where | What it does |
|---|---|---|
| Approvals | bridge / Redis | hold outbound messages for a human reaction |
| Egress allowlist | secrets proxy | restrict HTTP methods and destinations, inject credentials |
| Sheepdog | seccomp sandbox | enforce file/network/exec rules at the syscall level |
This page summarizes how they are configured. There is no runtime rule engine
or FilterAction-style API in the harness — governance is the sum of the
mechanisms below. Command, path, and syscall enforcement is done by sheepdog at
the kernel boundary (see below), not by a separate string-matching filter layer.
Approvals (human in the loop)
For outbound messages that should not be sent autonomously, the bridge approval protocol holds a message until a human reacts to a notification. Configure it per bridge:
services.cowboy.bridges.discord.approval = {
required = true;
notify = "discord";
notify_channel = "<channel-id>";
timeout = 3600;
};
Approval state is tracked in Redis hashes and resolved from human reactions. See Approvals & Outbox for the full protocol.
Egress allowlist (secrets proxy)
In the managed proxy deployment, network rules route HTTP through the proxy and the agent uses placeholder provider credentials. Local and ordinary Docker launches hold credentials instead. With egress control enabled, GET, HEAD, OPTIONS and TRACE pass the method gate; every other method needs an allowed destination or receives HTTP 403. Allowed requests can still change external state or disclose readable data. See Security Model.
services.cowboy.secretsProxy = {
enable = true;
domainMappings = {
"api.anthropic.com" = {
secretPath = "/run/agenix/anthropic-key";
headerName = "x-api-key";
};
};
# Extra write-allowed domains. Domains in domainMappings are implicitly
# write-allowed.
allowedWriteDomains = [ "github.com" "api.github.com" "*.githubusercontent.com" ];
};
The allowlist enforcement lives in the mitmproxy addon (proxy/addon.py);
write methods, the allowed-domain check, and wildcard matching are implemented
there. See Security Model.
Sheepdog (seccomp sandbox)
Sheepdog enforces file, network, and exec rules at the syscall level rather than
by string matching. Rules are verb-granular — Bash, Read, Edit, Create,
Delete, and Connect — and resolve to allow or deny, with optional
runtime-granted exceptions (lazy permissions) taking precedence over baked-in
denies.
services.cowboy.sheepdog = {
enable = true;
lazyPerms = true; # allow runtime permission grants
};
services.cowboy.agents.<name>.sheepdog = {
deny = [ "Connect(0.0.0.0/0)" ];
allow = [ "Read(/home/*/workspace/**)" "Edit(/home/*/workspace/**)" ];
blockedSyscalls = [ /* ... */ ];
readonlyPaths = [ /* ... */ ];
maskedPaths = [ /* ... */ ];
};
Sheepdog is Linux-only. See crates/sheepdog/src/policy.rs and
modules/options/sheepdog.nix.
See also
Plan state is advisory
The agent can call update_plan_state to record an objective, tasks, blockers
and the next step. Each task is pending, in progress or completed. Updates
replace the supplied fields, keep at most one task in progress and record an
update timestamp.
The state is saved as plan_state.json in the session directory and loaded
on resume. A non-empty plan appears in the model’s transient context packet,
including up to eight tasks. It does not need to be repeated in the transcript.
This is a memory aid, not an approval gate. Plan state does not restrict the tool list, impose a read-only planning phase or enforce that actions follow the listed tasks. Use bridge approvals and host policy for enforced controls.