Cowboy
Persistent agents on infrastructure you control.
Cowboy runs persistent AI agents on infrastructure you control, using Nix to configure their tools, credentials, message access, and approved actions.
The strongest documented deployment is managed NixOS on Linux/x86-64. Local and ordinary Docker paths have different boundaries and hold credentials locally. The OCI runtime is a specialized optional route for Cowboy’s own payload.
Why run an agent this way?
A maintenance task can produce a reviewed change without giving the agent the credential that publishes it. The repository’s effects runner check exercises adding a documentation post to a Git repository: it describes a commit bundle, applies it, and verifies the remote branch. This is a recorded test workflow, not a claim about production adoption.
In a deployment with the effects bridge and Discord approvals configured, the agent requests publication. The publishing runner checks the bundle and produces a card naming the target, execution identity, exact SHA, changed files, and diffstat. An authorized human reacts ✅ to push exactly that commit or ❌ to reject it. The runner executes under the configured publishing identity, checks policy again, and refuses a branch that moved after review. The result appears in the agent’s effects inbox and the configured operator status channel. Gated effects explains the configuration.
Choose a deployment
| Goal | Route | Boundary |
|---|---|---|
| Evaluate in a terminal | Local CLI and Zellij | Invoking user’s permissions; local credentials |
| Run a persistent managed service | NixOS systemd service running Montana | Per-agent component, shared service hardening envelope, network namespace, credential proxy, Sheepdog, and broker controls |
| Run on an ordinary Docker host | Self-contained image | Docker isolation; credentials in the container’s state volume |
| Integrate Cowboy’s payload with an OCI engine | Optional Cowboy OCI runtime | Bundle policy and configured cage; specialized command surface |
The full deployment matrix includes host requirements and state locations. “Production path” means the intended supported managed configuration, not measured adoption or general runtime certification. Its agent daemon runs Montana directly, without the OCI runtime.
Network restrictions, filesystem access, tool policy, and approval gates are separate controls. HTTP policy restricts methods and destinations; it does not promise that every permitted request leaves external state unchanged. Data an agent can read can leave through permitted requests or replies. Read the security model before granting sensitive access.
Start here
Continue to Installation, which maintains release availability and source-build instructions. Then use First request and daily operation and Troubleshooting.
Configuration explains how to change the
agent. Architecture explains the implementation
and the cowboy:agent@0.2.0 component contract.
Deployment paths
Choose the boundary before installing. The strongest documented deployment is managed NixOS on Linux/x86-64. Local and ordinary Docker paths deliberately hold credentials; the optional OCI route runs Cowboy’s own payload. These paths do not form a ladder, and tool behavior depends on each path’s policy.
| Path | Host requirements | Process identity and controls | Credentials | Mutable state | Primary limitation |
|---|---|---|---|---|---|
| Local evaluation | Nix source build; Zellij for the UI; Linux or macOS with the dependencies | Invoking user; Zellij harness or native Montana | Environment, user/system key files, or readable agenix files | Local XDG data/config directories and workspace | File and shell tools have the user’s permissions; no managed proxy or service envelope |
| Managed NixOS | NixOS, Linux/x86-64; configured operator and provider secret | Configured agent user; Montana under systemd, shared hardening envelope, network namespace, credential proxy, Sheepdog, broker ACLs | Proxy-readable secret files outside the agent; agent receives placeholders | Agent home, configured writable mounts, broker/bridge state | Controls depend on configuration; readable data is not confidential from permitted egress |
| Ordinary Docker | Source-built image; Docker on a matching Linux image architecture | Container root; ordinary Docker isolation; wizard uses Zellij, settings-based service uses Montana | Mode 0600 key file in the configuration volume | Configuration and workspace volumes; explicitly persist the data directory | No managed credential proxy or Cowboy cage; settings-based CLI launch does not itself mount the default data directory |
| Optional OCI cage | Linux/x86-64, Nix-built seed/rootfs, registered Cowboy runtime and resolver, Docker, companion images | Bundle-selected uid/gid; Cowboy runtime runs linked Montana with bundle syscall/filesystem policy | OpenRouter credential in companion; placeholder in agent | Bundle-selected writable mounts; companion scratch on host | Current CLI companion supports OpenRouter only; Redis companion has no durable volume; not a general container runtime |
Managed NixOS
Nix builds a per-agent component and declares the daemon, tools, proxy, message access, and service controls. The daemon launches Montana directly under systemd. It joins the configured Cowboy network namespace; the shared service envelope limits its writable world to the agent home and declared mounts. Sheepdog mediates tool execution. The daemon requires the credential proxy.
The module can also build an OCI bundle. That bundle is a separate deployment artifact, not an intermediate step in the systemd daemon launch. “Production” here describes the intended supported configuration, not adoption or runtime certification.
Local and ordinary Docker
Local UI sessions use the Zellij harness. Local headless sessions use Montana. Changing hosts does not add the managed security controls.
The ordinary image has two startup paths. The interactive wizard launches Zellij. A settings-based first boot generates and builds a Nix configuration, then its supervisor launches Montana on a Unix socket. Both keep real keys inside the container. The Docker walkthrough sets an explicit persistent data directory for the interactive path.
The NixOS module’s legacy container supervisor is a separate development artifact. It must not be confused with the managed systemd service or with the ordinary image’s settings-based supervisor.
OCI and embedding
The OCI runtime accepts Cowboy payloads and a narrow lifecycle command surface. It does not run arbitrary image commands. See the OCI contract before integrating it with an engine.
A custom host can implement cowboy:agent@0.2.0, but must supply and authorize
its own effects and filesystem access. Component portability alone supplies
no host security policy. Bare, user, and container are design
classifications, not three runtime selectors.
Continue to Installation. The security model describes the conditions behind the controls in this table.
Installation
Choose a route in the deployment matrix first. Local, managed NixOS, ordinary Docker, and the optional OCI cage have different credential and process boundaries. Policy can change which tools work.
Release availability
Nothing has been published to a package registry. There is no published Python package or prebuilt Cowboy image. This page is the maintained release availability statement; all installation routes below build from source.
Start with a checkout and Nix with flakes enabled:
git clone https://github.com/dmadisetti/cowboy.git
cd cowboy
Keep the revision and lock file used for a deployment. The commands below run from this checkout unless they explicitly refer to your host configuration.
Local evaluation
Build the CLI and the repository’s Zellij variant:
nix build .#get-cowboy --out-link /tmp/cowboy-cli
nix build --impure --expr '
let f = builtins.getFlake (toString ./.);
in (import f.inputs.nixpkgs {
system = builtins.currentSystem;
overlays = [ f.overlays.zellij ];
}).zellij' --out-link /tmp/cowboy-zellij
export PATH=/tmp/cowboy-cli/bin:/tmp/cowboy-zellij/bin:$PATH
read -rsp 'Anthropic API key: ' ANTHROPIC_API_KEY
export ANTHROPIC_API_KEY
cowboy --model anthropic:claude-opus-4-5-20251101
The Zellij overlay supplies the memory limit used by the harness. The key is in the launching environment. The CLI packages its harness WASM; a plain Python install from the checkout does not stage that artifact.
In Zellij, send:
Read README.md and summarize what this project does. Do not change files.
Expect tool activity and a model reply in the session, not a predetermined answer. File and shell tools run with your permissions. From another shell in the same environment, stop it with:
cowboy stop cowboy
For a local headless evaluation, build Montana and the component explicitly:
nix build .#montana --out-link /tmp/cowboy-montana
nix build .#agent-component-lite --out-link /tmp/cowboy-component
export MONTANA_BIN=/tmp/cowboy-montana/bin/montana
export MONTANA_WASM=/tmp/cowboy-component/share/cowboy/cowboy-agent.wasm
cowboy serve --model anthropic:claude-opus-4-5-20251101
This prints a Unix socket connection command. Connect to that socket and send one JSON line:
{"submit":"Say hello and describe your available tools."}
The socket emits display frames containing activity and the answer. Stop the
unnamed headless instance with cowboy stop agent. This path still holds real
keys and runs with your permissions; headless operation adds no confinement.
Managed NixOS
Use an existing Linux/x86-64 NixOS host configuration. Add Cowboy as a flake
input, import the Cowboy and Home Manager modules into that host’s module
list, and keep its lock file under version control. The fragment below assumes the host
already imports agenix and has an operator account named alice; substitute
your operator name. The encrypted key must be decryptable by this host.
# Add to the inputs of your host flake:
inputs.cowboy.url = "github:dmadisetti/cowboy";
# Include in the host's nixosSystem modules list:
modules = [
./configuration.nix
cowboy.nixosModules.default
cowboy.inputs.home-manager.nixosModules.home-manager
./cowboy-agent.nix
];
Save this module as cowboy-agent.nix beside the host flake:
{ lib, ... }:
{
programs.fish.enable = true;
services.cowboy.broker = "alice";
services.cowboy.agents.dev = {
enable = true;
user = "dev";
uid = 1338; # Choose an unused UID.
model = "anthropic:claude-opus-4-5-20251101";
daemon.enable = true;
daemon.expose = 4312;
};
age.secrets.anthropic-key = {
file = ./secrets/anthropic-key.age;
owner = "cowboy-proxy";
};
services.cowboy.secretsProxy.domainMappings = lib.mkForce {
"api.anthropic.com" = {
secretPath = "/run/agenix/anthropic-key";
headerName = "x-api-key";
};
};
}
The encrypted file is part of your host configuration; the plaintext key never belongs in Nix text or the store. Its runtime reference is the proxy mapping’s secret path. Restricting the mapping to this provider avoids provisioning unneeded provider and search credentials.
Build and activate from your host configuration directory, replacing HOST with its NixOS configuration name:
sudo nixos-rebuild switch --flake .#HOST
systemctl status cowboy-serve-dev --no-pager
cowboy connect 127.0.0.1:4312#/var/lib/cowboy/expose/dev.token
Run the connection command as the configured operator. It opens the Zellij
client to the managed agent. Send “Say hello and describe your available
tools.” Expect activity and a reply there; inspect service diagnostics with
journalctl -u cowboy-serve-dev. The localhost control token grants control,
including the ability to answer approval prompts. A read-only observer needs
the separate observer endpoint and token.
The daemon runs Montana directly under systemd as dev, with its per-agent
component, namespace, proxy, Sheepdog, and shared service envelope. It does not
invoke the OCI runtime. To leave the daemon stopped:
sudo systemctl stop cowboy-serve-dev
A socket quit is insufficient: the unit restarts automatically. Set
daemon.autostart = false and rebuild if it should remain off at boot.
Docker
Build on a Nix machine matching the target image architecture, then load the archive on the Docker host. The target host needs Docker, not Nix:
nix build .#docker-image --out-link /tmp/cowboy-image
docker load < /tmp/cowboy-image
docker run -it --name cowboy-eval \
-e XDG_DATA_HOME=/etc/cowboy/data \
-v cowboy-etc:/etc/cowboy \
-v cowboy-workspace:/opt/workspace \
cowboy:latest
For separate machines, transfer the image archive before loading it. On first
boot the wizard asks for provider, model, and key. It writes the key with mode
0600 into the configuration volume, then opens the setup agent in Zellij.
Tell it what work you want help with, your communication preferences, and when
it should ask before acting. It saves the agreed profile to
/opt/workspace/.cowboy/prompt.md; an otherwise empty workspace is normal.
Choosing a role such as calendar assistant does not authorize or configure a
calendar integration. Setup should state which capabilities work and which
still need access, while preserving existing workspace files.
The explicit data directory keeps session state in the named volume, including
on later starts. Tools run as container root with ordinary Docker isolation;
there is no managed credential proxy or Cowboy cage.
When the profile is saved, stop and resume from the Docker host. The normal agent loads that profile and resumes the conversation from persistent state:
docker stop cowboy-eval
docker start -ai cowboy-eval
Without a TTY or staged settings, a fresh image exits with “no configuration found.” For a settings-based headless service, the CLI can stage a bootstrap settings file and build a per-agent generation. That route is distinct from the interactive wizard. Its current launcher mounts config and workspace but does not set a persistent Montana data directory; account for that before relying on it for durable sessions. See daily operation.
Optional OCI cage
This specialized route requires Linux/x86-64, Docker, a source checkout, and a
registered Cowboy runtime with its imageless resolver. Follow the registration
instructions in crates/runtime/smoke.sh and the
OCI contract.
Registration changes Docker’s daemon configuration; it is not supplied by
installing the CLI.
Once the runtime is registered, prepare the Cowboy payload and the placeholder image, then launch from the checkout with the CLI available:
nix build .#cowboy-agent-seed --out-link /tmp/cowboy-seed
nix build .#cowboy-agent-rootfs --no-link
tar -C /tmp/cowboy-seed/rootfs -c . | docker import - cowboy-imageless:latest
read -rsp 'OpenRouter API key: ' OPENROUTER_API_KEY
export OPENROUTER_API_KEY
cowboy serve demo --sandbox cage --expose 4210
The companion currently injects only OpenRouter credentials; another provider’s key will not work here. It holds the real key and exposes egress through Unix sockets. Docker runs the payload with the Cowboy runtime and networking disabled; the runtime launches linked Montana under the bundle’s policy.
Use the printed cowboy connect command and send the same hello request.
Expect display frames in the client. The component and model come from the
seed configuration. Stop the cage and its companions with cowboy stop demo.
The companion Redis has no persistent volume; this is not the managed NixOS
service’s persistence contract.
Continue with First request and daily operation.
First request and daily operation
Use one of the complete routes in Installation. Start with a request whose result you can check, such as “Read README.md and summarize what this project does. Do not change files.” Check the tool results and answer in the client. A successful launch alone says nothing about model authentication or useful work.
Inventory, status, and health
cowboy list
cowboy status
cowboy doctor
These answer different questions. The list shows registered agents that can run. Status shows runtime instances and their last reported activity. Doctor checks whether services are up, the event loop responds, and inbox consumers are still reading. A local session need not have a registry entry, and doctor cannot verify every deployment from every account. An unknown result is not a pass. Use Troubleshooting to interpret it.
For a named managed agent:
cowboy status dev
cowboy doctor dev --json
journalctl -u cowboy-serve-dev --since today
cowboy logs reads the launch log of a backgrounded local serve process; it is
not a general transcript or journal reader. For example, cowboy logs agent
reads the unnamed headless instance’s log when it was started with --daemon.
Use the system journal for managed units, Docker logs for containers, and the
client/session history for conversation content.
Review external actions
Approvals are configured per bridge or effect. For a gated Git push, review the runner’s target, execution identity, commit SHA, changed files, and diffstat. The agent’s prose is not the runner’s description. An authorized approver chooses ✅ to push exactly that SHA or ❌ to reject it. An expired request needs a new review; a branch that moved requires a newly prepared request.
The operation result returns to the agent’s inbox and the configured status channel. Keep the request ID and SHA if delivery or execution is uncertain. An accepted message or approval is not proof that the push or rebuild finished. See approvals and gated effects.
Restart and recover
Managed daemons restart after exit, including clean socket quits, with a five-second delay and a start-rate limit. A changed per-agent component also triggers a restart on NixOS activation. Use systemd for an explicit restart:
sudo systemctl restart cowboy-serve-dev
cowboy doctor dev
A deliberate systemd stop leaves the daemon down. Ordinary containers in the
installation example have no Docker restart policy; start them explicitly.
For local headless and cage launches, retain the original launch command:
cowboy restart cannot reconstruct their transient socket, sandbox, and
exposure flags. It stops those instances and asks you to relaunch. For a
settings-based Docker agent it can reuse the declared volume.
After a crash, a sender may receive no answer even though a request was accepted, or see a duplicate after redelivery. Resumed work can also lose the original reply destination; inspect state and external outcomes before resubmitting. Message reliability gives the exact recovery contract.
Back up mutable state
Stop the agent and its writers before copying state. Keep ownership and modes, especially for keys. A Nix closure rebuilds software; it does not restore conversations, working files, credentials, or pending approvals.
- Local: preserve the workspace and Cowboy’s XDG data/config directories
(normally under
~/.local/shareand~/.config). - Managed: preserve the configured agent home and writable mounts, plus the enabled broker/bridge services’ persistent state and secret sources. Include Redis persistence when recovering queued messages and approvals matters.
- Docker walkthrough: preserve both named volumes, including the explicit data directory under the configuration volume. Container removal does not remove these named volumes.
- Settings-based Docker CLI: config and workspace are under
~/.local/share/cowboy/agents/<name>/by default. The current supervisor leaves Montana’s default data directory in the container layer. Preserve it separately before replacing the container, or configure a data directory in a mounted volume in the generated supervisor. The config/workspace mounts alone are not a complete session backup. - OCI: preserve the writable mounts you declared in the bundle; do not treat runtime scratch or the CLI companion’s Redis as a durable backup.
Upgrade and rebuild
Keep the old revision, lock file, and a state backup until the new deployment has answered a test request. Build the new CLI or image from the chosen source revision. On NixOS, update the host’s Cowboy input and rebuild the host, then check doctor and the first real request.
For the Docker walkthrough, stop and remove the old container after backing up,
load the new image, and repeat the run command with the same named volumes. A
settings-based container also has a generated flake and result link under its
configuration volume. Edit that persistent configuration and run cowboy bake
inside the container to build a generation; restarting with an existing result
link does not itself rebuild it. A new base image alone does not replace that
existing generation.
An approval-gated rebuild pins the main commit and resolved input overrides before asking for approval. The result reports when a successful build produced the same system or home generation. Treat that as no deployment change, not evidence that the requested behavior reached the running agent. See rebuild diagnostics.
Stop and remove
Stop local or CLI-managed agents with cowboy stop <name>. For a managed
daemon, stop its systemd unit, disable the agent in the NixOS configuration,
and rebuild. For the Docker walkthrough, stop and remove the named container.
Keep backups, then deliberately remove only the state and volumes you no
longer need. Removing software or disabling a service is separate from erasing
its history and keys. Revoke provider or publishing credentials when retiring
the deployment.
Continue with Configuration to change models, tools, message access, and service settings.
Troubleshooting
Start with cowboy doctor. A service can be up while the agent has stopped
reading messages.
cowboy list
cowboy status
cowboy doctor
cowboy doctor dev --json
The list is the registered inventory. Status observes running instances and last reported activity. Doctor separates three observations:
| Observation | What it establishes | What it does not establish |
|---|---|---|
| Process/service exists | The host still has a running process | The event loop is making progress |
| Event loop responds | A diagnostic request produced a new frame | Inbox polling still runs |
| Inbox consumer is active | Redis records recent reads | The model finished a turn or a reply reached its destination |
A replayed status frame can outlive a stalled loop. Doctor asks for a new one. Its attention check examines inbox consumer activity independently.
Checks report ok, fail, or unknown. Unknown means the probe could not establish an answer, for example because the caller cannot read an agent’s Redis socket or the deployment has no supported probe. Read the reason; do not translate it into healthy. On a managed host, rerun as an authorized operator with the required privilege if the reason is access. Do not broaden socket permissions to make a diagnostic green.
Find the right log
# Managed service:
systemctl status cowboy-serve-dev --no-pager
journalctl -u cowboy-serve-dev --since today
systemctl list-units 'cowboy-*'
# Backgrounded local serve (default headless name):
cowboy logs agent
# Docker walkthrough:
docker logs cowboy-eval
Inspect the first failed service and its dependencies. The managed daemon requires the namespace and credential proxy; it does not launch an OCI bundle. A cage startup failure instead belongs to Docker, its runtime registration, the seed/rootfs, bundle policy, or companion services.
The CLI or Zellij is missing
Use the source-build installation. Select the CLI package explicitly; the default flake package is not the Python launcher. A Python source install alone does not bundle the harness WASM. Zellij is required for interactive sessions and the connect client, not for Montana’s headless service.
The client opens but the model fails
Check the provider/model specification, the credential source for this route, and the upstream response. Local and ordinary Docker paths hold real keys; managed agents use placeholders and need the proxy’s secret path to be readable by its service user. User and system keys files accessible to group or others are ignored. See secret discovery.
On a managed daemon, a runtime provider switch also needs provider access in
the Nix configuration. A proxy credential alone does not supply the component’s
provider key placeholder. Declare the provider through a model option or
extraProviders and rebuild.
Record the provider, model, HTTP status, and request ID without copying keys. A method or destination denial calls for reviewing the configured policy, not bypassing the proxy.
A tool is denied
Identify the layer from its result: tool approval, filesystem permissions, Sheepdog syscall policy, service hardening, proxy methods/destinations, or OCI policy on the cage route. Compare the operation to that deployment’s declared access. A successful local run is not evidence that the managed policy should permit it.
An approval or reply seems stuck
Follow the request through ingest, the agent inbox, the agent turn, and the outbox. For Discord, the receipt reaction establishes that the bridge queued the message, not that the agent answered. Check doctor for continued inbox consumption and inspect whether an approval is pending, rejected, expired, or superseded. Only configured approvers can resolve a gated Discord request.
For Git effects, preserve the request ID and SHA. The approval must match the runner’s description and the applied SHA. A branch moving after review causes a refusal; fetch and prepare a fresh request instead of treating the old approval as permission for another commit. A push can succeed while its verification fetch fails: inspect the remote before retrying.
After interrupted work, a request may have no reply, be redelivered, or resume without its original reply destination. Inspect external results before asking for the same side effect again. See message reliability.
A rebuild succeeds but nothing changes
The rebuild service compares the built system closure with the running system, or the new home activation path with the prior one when it can read that state. An identical result gets an explicit no-change note. A successful command can therefore mean no new deployment.
First check the rebuild type: a home rebuild cannot deploy system services, networking, or system packages. Check that the approved commit contains the change, that it reached the tracked branch, and that the flake imports the edited file. Check resolved input overrides in the approval and result. Then check the deployed generation and whether the daemon restarted onto the new component. Doctor compares the generated agent configuration timestamp with the daemon start time to flag possible configuration drift. If the rebuild could not establish a prior home generation, absence of a no-change note is not proof of change.
For settings-based containers, an existing result link is reused at startup.
Build the edited generated configuration with cowboy bake inside the container
before restarting; merely editing JSON or replacing the base image does not
necessarily replace that generation.
Restart without losing the evidence
Copy mutable state before repair. Managed daemons restart after exit, but a live stalled process need not exit; record doctor and journal output before an explicit systemd restart. Recheck health and send a small request afterward. See daily operation for restart limits, backup, and removal.
Report a useful bug
Include the source revision, OS/architecture, deployment route, redacted command or Nix configuration, expected and observed behavior, doctor output, and a bounded log excerpt. For bridge/effect problems include the request ID and commit SHA. Report at GitHub issues.
Configuration
CLI launch settings, persistent JSON, generated NixOS configuration, and secret discovery are different inputs. A provider key is not another level in ordinary settings precedence.
Local settings
The launcher merges JSON objects in this order, with later files winning:
/etc/cowboy/config.json~/.config/cowboy/config.json(under the configured XDG config directory).cowboy/config.jsonin the current directory
Missing, unreadable, malformed, and non-object files contribute no settings. Explicit CLI options override the corresponding ordinary settings.
For provider, main model, and heartbeat, merged files override generated agent JSON, which overrides built-in defaults. For secondary models and memory/rerank settings, generated agent JSON overrides merged files; supported CLI options still win. Do not assume one precedence order covers every field. The declared systemd daemon takes its arguments from NixOS configuration directly, rather than reading the operator’s local launch defaults.
A local config can contain:
{
"provider": "anthropic",
"model": "claude-opus-4-5-20251101",
"log_level": "info"
}
Common settings subset
This is a common subset, not the entire recognized surface. The implementation
inventory is crates/contracts/src/keys.rs; a host or launcher need not expose
every component key as a CLI option.
| Key | Meaning |
|---|---|
provider | Provider for the main model |
model | Bare model ID in local JSON; provider comes from the separate key |
summary_model | Provider:model specification for summarization |
compact_model | Provider:model specification for compaction |
subagent_model | Provider:model specification for sub-agents |
vision_model | Native vision or a separate provider:model specification |
memory_backend | Memory search backend |
heartbeat | Launcher heartbeat interval, in seconds |
log_level | Logging verbosity |
The component key for heartbeat is heartbeat_interval; the launcher’s JSON
setting and flag use heartbeat and --heartbeat. Model availability and
provider access are separate: selecting a model does not provision credentials.
Use cowboy models to inspect the bundled catalog and each command’s --help
for its options.
Secrets
Outside proxy deployments, discovery checks each key in this order and uses the first nonempty value:
- Provider environment variable, such as
ANTHROPIC_API_KEY. - User keys file,
~/.config/cowboy/keys.json. - System keys file,
/etc/cowboy/keys.json. - The provider’s agenix file, if readable.
Keys files map component key names, such as anthropic_api_key, to secret
values. They must have mode 0600: files accessible to group or other users are
skipped with a warning. Secret files and environment variables remain real
credentials in local and ordinary Docker deployments.
Managed daemons receive proxy placeholders. The actual key belongs at the secret path named by the proxy domain mapping, readable by the proxy service user. The managed installation shows an agenix declaration with that ownership.
First-boot settings
A --settings file for cowboy init or cowboy serve is a validated bootstrap
input, not the general config file. It accepts a key reference through
api_key_env or api_key_file; a literal api_key is rejected. For example:
{
"agent_name": "myagent",
"posture": "container",
"provider": "anthropic",
"model": "claude-opus-4-5-20251101",
"api_key_env": "ANTHROPIC_API_KEY"
}
The host resolves the reference and writes the key into the agent’s configuration volume with mode 0600. A bootstrapped volume rejects restaging; edit its persistent configuration instead. See daily operation for generation and backup considerations.
NixOS module
Import the default NixOS module and enable individual agents under
services.cowboy.agents. There is no top-level service enable switch. Keep
agent usernames equal to their agent names when using Redis ACLs, and give
each enabled agent a unique UID.
The main model option takes a provider:model specification. Common per-agent
options include homeDirectory, prompts.system, mounts, daemon.enable,
and daemon.expose. The latter provides an authenticated control endpoint;
daemon.observe provides a separate read-only endpoint and credential.
Nix generates per-agent launch configuration and components, and message-source
configuration for agents and bridges. Edit the Nix declaration, not the
generated files under /etc/cowboy. The full per-agent schema is in
modules/options/user.nix.
Managed provider access is provisioned from the model options. To keep another
keyed provider available for a runtime model switch, declare it in
extraProviders and provision its proxy secret. This grants provider access;
it does not select the main model.
Continue with tools and skills and bridge integration.
Approvals & Outbox
title: Outbox Approval Protocol tags: [pubsub, approval, discord, email]
Outbox Approval Protocol
Outbox services can hold an outbound message for human approval before sending it. The agent never sends directly; it writes to an outbox stream, and a broker service decides whether to forward the message, optionally gating it behind a human reaction.
Source: pkgs/bridge/base.py (OutboxService), pkgs/bridge/ingest.py
(IngestService), with per-platform implementations in pkgs/discord/ and
pkgs/email/.
See also Security Model.
Transport
Messages move over Redis Streams. Each bridge {name} has an inbox stream
({name}:inbox) and an outbox stream ({name}:outbox). OutboxService reads
its outbox via a consumer group ({name}-outbox / worker {name}-worker) and
acknowledges each entry after handling it. Approval state lives in plain Redis
hashes so any service can poll it.
Flow
When approval_required is set, an outbound message is held and routed through
a notification channel:
- The agent writes a message to its outbox —
{agent}:{source}:outboxunder the per-agent ACL (the flat{source}:outboxon a shared Redis). OutboxService.process_one()seesapproval_requiredand the message has noapproval_id, so it parks the entry (_resume_or_start_approval):- creates
approval:{uuid}withstatus=pending,source,channel_id, acontent_preview(first 200 chars),created_atandexpires_at - writes a notification to the configured notify outbox
(
{approval_notify}:outbox) carrying theapproval_id - returns without acking, so the rest of the stream keeps moving
- creates
- The notify service (e.g. Discord) sends the notification. Because the message
carries an
approval_idand the send returns anexternal_id, the base loop records the mapping inapproval:{source}_map(external id → approval id). - A human reacts to the notification.
- The notify service’s ingest handler calls
OutboxService.resolve_approval(), which looks up the approval id from the map and, if the record is stillpendingand inside its window, setsstatustoapprovedorrejectedplus theapprover. A late answer — a record already decided, pastexpires_at, or already settled and deleted — is ignored; the map entry is dropped either way. - The original outbox service checks its parked entries once per loop
(
poll_awaiting, viacheck_approval, which reads the deadline before the decision); it sends the message onapproved, or acks it and tells the agent on its inbox that it wasrejectedorexpired.
State machine
approval:{uuid} = {
status: pending | approved | rejected | expired
source: email | discord | ...
channel_id: destination (email address, channel id, ...)
content_preview: first 200 chars of the message body
created_at: unix timestamp
expires_at: unix timestamp; past it the record reads as expired
approver: who resolved it (set on resolution)
}
pending --+-- approve --> approved --> send
+-- reject --> rejected --> drop, notify the agent
+-- timeout --> expired --> drop, notify the agent
expired is never written: it is what check_approval answers for a record
past expires_at, whatever its status field says. poll_awaiting deletes
the approval hash when it settles the entry.
Redis keys
approval:{uuid} # HASH — approval state
approval:{source}_map # HASH — external message id -> approval id
The requesting service writes and polls approval:{uuid}. The notify service
owns approval:{source}_map. Bridges share only the key format — Discord and
email do not know about each other.
Implementation
The park-and-poll flow (process_one(), poll_awaiting()), the static
resolve_approval(), and the auto-tracking of the map all live in the shared
OutboxService. A per-platform bridge only needs to:
- Return
SendResult(ok=True, external_id=...)from itssend(). - Call
OutboxService.resolve_approval(pubsub, source, external_id, status, approver)from its ingest handler when a reaction or reply arrives (IngestService.try_resolve_approval()wraps this).
systemd services
Each bridge {name} is run as three systemd services: cowboy-{name}-ingest,
cowboy-{name}-outbox, and cowboy-{name}-ping.
Configuration
Approval is configured per bridge under
services.cowboy.bridges.<name>.approval (and equivalently on
services.cowboy.pubsub.sources.<name>.approval):
services.cowboy.bridges.discord.approval = {
required = false; # hold outbound messages for manual approval
notify = "discord"; # which outbox to send approval notifications to
notify_channel = ""; # channel id within that outbox
timeout = 3600; # auto-reject after N seconds (0 = no timeout)
};
If required is true, notify_channel must be set. The Redis ACL enforcement
option keeps the agent restricted to stream commands so it cannot touch the
approval hashes directly.
Two bridges use this protocol for something other than a message. rebuild
gates a system rebuild, and effects gates a declared action on the world
(a git push, for instance) — see Gated Effects. For effects,
approval.required is read-only true: there is no unapproved mode, and the
approval is bound to a specific commit sha rather than to the request text,
so what the operator approved is what runs, byte for byte.
Adding a new approval channel
To approve via something other than Discord:
- Have the new service’s outbox return
external_idfromsend(). - Have its ingest call
OutboxService.resolve_approval(...). - Set
approval.notify = "<service>"on the bridge that needs approval.
No changes to OutboxService or the requesting service are required.
Gated Effects
An effect is an action on the world that an agent may propose but must not
perform: pushing to a branch someone else consumes, publishing, deploying.
services.cowboy.pubsub.effects gives each declared effect a fixed shape —
the agent names a commit, a human sees exactly that commit on an approval
card, and after ✅ the host applies exactly that commit as a user of the
operator’s choosing. Nothing the agent says on the card is trusted; nothing
the agent can do between the card and the push changes what gets pushed.
It reuses the Outbox Approval Protocol unchanged. What is new
is the seam on the far side of the gate: a bridge with no privilege of its own
starts a static systemd template unit through polkit, and that unit — running
as the declared runAs user inside the standard sandbox — does the work.
The flow
agent effects bridge host
----- -------------- ----
commit in workspace
effects-request tool ───▶ effects:outbox entry
(git bundle into pre_send:
the handoff dir) systemctl start --wait
cowboy-effect-<name>@describe-<sha> ──▶ unit (User=runAs):
unbundle, fetch, policy,
write results/describe-<sha>.json
read the result; refuse ⇒ ack + message
ok ⇒ pin sha/base/description
approval card → notify channel
⏳ … human reacts ✅
send():
guard pinned sha == requested sha
mark_effect
systemctl start --wait
cowboy-effect-<name>@apply-<sha> ─────▶ unit: re-fetch, re-run policy,
push --force-with-lease,
verify, write apply-<sha>.json
result → agent inbox [effects <id>] …
→ notify channel (untagged copy)
Two runs of the same program, addressed by the same 40-hex sha: describe before the card, apply after it. The sha in the unit instance name is the whole of what the bridge can choose; polkit permits only that one action, on only those unit names, for only the bridge user.
Declaring an effect
services.cowboy.pubsub.effects = {
enable = true;
approval.notify_channel = "<discord channel id>"; # required; there is no unapproved mode
declared.blog = {
runAs = "<user>"; # the unit's User=; never root, never the bridge
gitPush = {
remoteUrl = "git@github.com:me/blog.git"; # ssh://, git@, or file:/// — never https
branch = "main";
webUrl = "https://github.com/me/blog"; # optional; result messages link the commit
allowedPaths = [ "docs/posts/" ]; # directories (or exact files) every commit must stay inside
maxCommits = 3;
knownHosts = "github.com ssh-ed25519 AAAA…"; # required for ssh remotes
};
credentials.sshkey = "/home/<user>/.ssh/keys/github"; # loaded with LoadCredential=, readable by root only
};
};
approval.required is read-only true. An effects bridge that starts without
approval configured exits at authenticate(), and the module refuses to
evaluate without a notify channel or with the notify bridge disabled — the
failure mode where a “gated” effect quietly evaluates to an ungated one is the
one this module exists to close.
Each declared effect produces:
- a group
cowboy-effect-<name>whose members are the agents plusrunAs(never the bridge), and a setgid handoff directory/var/lib/cowboy-handoff/<name>(2770) where the agent’s bundles land; - a state directory
/var/lib/cowboy-effect-<name>owned byrunAsholding the mirrorrepo.git(0700),results/, and the lock; - a template unit
cowboy-effect-<name>@.servicewithUser=runAs,NoNewPrivileges,ProtectSystem=strict,ProtectHome, an empty capability set,IPAddressDenyoutside the remote, and the docker socket inaccessible; - one polkit rule allowing
cowboy-bridge-effectstostart(and only start)cowboy-effect-<name>@(describe|apply)-<40 hex>.service; - an
effects-requesttool visible to the effect’s agents.
Set runner instead of gitPush to supply your own program. It receives the
instance name (describe-<sha> or apply-<sha>) as $1, runs as runAs
with STATE_DIRECTORY set, and must write
$STATE_DIRECTORY/results/<verb>-<sha>.json with at least {"ok": bool, "message": str}; a describe result should also carry what the card shows
(base, count, commits, files, diffstat, policy).
What the agent does
The effects-request tool takes effect, sha, repo, and content. It
verifies the sha is a commit in repo and that refs/remotes/origin/<branch>
exists there, writes a git bundle of sha ^origin/<branch> into the handoff
directory (size-capped), and files one entry on effects:outbox. It returns
{"status": "pending_approval", "request_id": …} and nothing has run.
The outcome arrives later in the agent’s inbox as a message starting
[effects <request_id>]. A refusal — the branch moved, a file outside
allowedPaths, a merge commit, too many commits, a symlink — is text, not a
retry: the agent rebases and files a new sha, which is a new approval. The
policy is checked on every commit in the range, not on the net diff, so an
add-then-revert pair cannot smuggle a path into history that the card never
showed. A newer request for the same effect from the same agent supersedes
an older one still waiting; a refused request supersedes nothing.
What the human sees
Approval needed [effects/blog]: push to git@github.com:me/blog.git main (as <user>)
Requested by agent: `publish monix — reviewed, no blockers`
Commit 3f9c2a1e… (fast-forward from 096ab9d, 1 commit, policy ok)
3f9c2a1 blog: publish monix (agent-gated) (agent, 2026-08-25)
Files:
A docs/posts/2026-08-25-monix.md
1 file changed, 118 insertions(+)
Inspect: git -C /var/lib/cowboy-effect-blog/repo.git show 3f9c2a1
React ✅ to push exactly this commit, ❌ to reject. Expires in 60 min; if the branch moves first the push refuses itself.
Every line but the quoted Requested by one comes from the describe run,
which read the bundle as runAs, not from the agent. The agent’s own text
is a single quoted line, capped, with backticks neutralised, so it cannot
imitate the trusted lines. The Inspect: path is a real mirror on the host
for a second look before reacting.
What can still go wrong
- The branch moves.
applyrefetches and requires the branch to be at the basedescriberecorded; the push itself is--force-with-leaseagainst that base. Either check failing is a refusal with the new head named. - The unit is interrupted. The bridge marks the entry
phase=effectbefore startingapply, so a restart mid-push reports uncertainty instead of re-running.on_uncertainreadsresults/applied-<sha>.jsonand says whether the push in fact completed. - A stale result. The bridge treats a result file older than the unit it just started as absent — a describe that failed to run (polkit said no, say) cannot be answered by the previous run’s file.
runAsis a login user. The unit inherits that user’s DB groups (the module warns). The sandbox is what bounds them; a dedicated uid with a deploy key is the better end state, and is a one-line change torunAsandcredentials.sshkey.
See Approvals & Outbox for the protocol underneath and
specs/THREAT-MODEL.md §8 for the residual authorities this adds.
Rebuilds and generated environments
The rebuild bridge is separate from a declared Git effect. Enable it through
services.cowboy.pubsub.rebuild and configure its repository and approval
settings. The rebuild-request tool queues a system or home rebuild and returns
a request ID. A queued or pending-approval result is not a completed rebuild.
The later outcome arrives in the agent inbox tagged [rebuild <request_id>].
A system request builds the configured remote revision with nixos-rebuild boot, then starts activation in a separate systemd unit. A home request builds
and activates the requesting agent’s home-manager profile. A home rebuild
cannot apply system service or networking changes. Local unpushed edits are
not the remote revision the bridge builds.
For system requests, a success message means the build finished and activation was launched; it does not prove the asynchronous switch completed. Inspect the activation unit and running service before calling the change live:
sudo systemctl status cowboy-activate.service
sudo journalctl -u cowboy-activate.service
cowboy doctor
The bridge compares the built system with the previously running generation,
or the built home profile with the previous home generation. When it can prove
they are equal, its result says the rebuild deployed no change. Check the
request type, pushed revision and selected flake inputs before retrying. An
unreadable generation is unknown, not evidence of a no-op. After an interrupted
activation, inspect the running system and the bridge’s status record at
/var/lib/cowboy-rebuild/last-rebuild.json before submitting again.
Ordinary Docker also has implemented generation machinery: settings can scaffold
a flake in the persistent configuration volume, bake its container output and
launch the resulting supervisor. The host records generation metadata for
cowboy list; that record is not a liveness check. Configuration and workspace
volumes are mutable state to preserve across replacement. See
Deployment Paths for launch setup.
These mechanisms generate and apply environments. They do not implement an autonomous rollback or pull-request review workflow.
AI Governance
Cowboy constrains what an agent can do through several independent mechanisms. They overlap deliberately: a command rejected by one layer is not relied upon to be caught by another.
| Layer | Where | What it does |
|---|---|---|
| Approvals | bridge / Redis | hold outbound messages for a human reaction |
| Egress allowlist | secrets proxy | restrict HTTP methods and destinations, inject credentials |
| Sheepdog | seccomp sandbox | enforce file/network/exec rules at the syscall level |
This page summarizes how they are configured. There is no runtime rule engine
or FilterAction-style API in the harness — governance is the sum of the
mechanisms below. Command, path, and syscall enforcement is done by sheepdog at
the kernel boundary (see below), not by a separate string-matching filter layer.
Approvals (human in the loop)
For outbound messages that should not be sent autonomously, the bridge approval protocol holds a message until a human reacts to a notification. Configure it per bridge:
services.cowboy.bridges.discord.approval = {
required = true;
notify = "discord";
notify_channel = "<channel-id>";
timeout = 3600;
};
Approval state is tracked in Redis hashes and resolved from human reactions. See Approvals & Outbox for the full protocol.
Egress allowlist (secrets proxy)
In the managed proxy deployment, network rules route HTTP through the proxy and the agent uses placeholder provider credentials. Local and ordinary Docker launches hold credentials instead. With egress control enabled, GET, HEAD, OPTIONS and TRACE pass the method gate; every other method needs an allowed destination or receives HTTP 403. Allowed requests can still change external state or disclose readable data. See Security Model.
services.cowboy.secretsProxy = {
enable = true;
domainMappings = {
"api.anthropic.com" = {
secretPath = "/run/agenix/anthropic-key";
headerName = "x-api-key";
};
};
# Extra write-allowed domains. Domains in domainMappings are implicitly
# write-allowed.
allowedWriteDomains = [ "github.com" "api.github.com" "*.githubusercontent.com" ];
};
The allowlist enforcement lives in the mitmproxy addon (proxy/addon.py);
write methods, the allowed-domain check, and wildcard matching are implemented
there. See Security Model.
Sheepdog (seccomp sandbox)
Sheepdog enforces file, network, and exec rules at the syscall level rather than
by string matching. Rules are verb-granular — Bash, Read, Edit, Create,
Delete, and Connect — and resolve to allow or deny, with optional
runtime-granted exceptions (lazy permissions) taking precedence over baked-in
denies.
services.cowboy.sheepdog = {
enable = true;
lazyPerms = true; # allow runtime permission grants
};
services.cowboy.agents.<name>.sheepdog = {
deny = [ "Connect(0.0.0.0/0)" ];
allow = [ "Read(/home/*/workspace/**)" "Edit(/home/*/workspace/**)" ];
blockedSyscalls = [ /* ... */ ];
readonlyPaths = [ /* ... */ ];
maskedPaths = [ /* ... */ ];
};
Sheepdog is Linux-only. See crates/sheepdog/src/policy.rs and
modules/options/sheepdog.nix.
See also
Plan state is advisory
The agent can call update_plan_state to record an objective, tasks, blockers
and the next step. Each task is pending, in progress or completed. Updates
replace the supplied fields, keep at most one task in progress and record an
update timestamp.
The state is saved as plan_state.json in the session directory and loaded
on resume. A non-empty plan appears in the model’s transient context packet,
including up to eight tasks. It does not need to be repeated in the transcript.
This is a memory aid, not an approval gate. Plan state does not restrict the tool list, impose a read-only planning phase or enforce that actions follow the listed tasks. Use bridge approvals and host policy for enforced controls.
Memory System
Persistent, cross-session memory for the agent. Notes are stored as Markdown files and surfaced back into the model’s context via a search step before the agent acts.
Source: crates/core/src/memory.rs. Tools are registered in
crates/core/src/tools/builtin.rs.
Backends
Memory is pluggable behind the MemoryBackend trait. Two backends exist:
ZkBackend— zettelkasten built on thezkCLI. Keyword-based full-text search. Always available, including lite builds. This is the default for lite.QmdBackend— hybrid BM25 + vector search via theqmdCLI, usingzkfor note authoring. Available only in ranch (non-lite) builds, where it is the default (memory_backend = "qmd").
The backend is selected by the memory_backend config key ("qmd" or "zk").
qmd_min_score (default 0.4) sets the relevance cutoff for the qmd backend.
If a lite build is asked for qmd, it falls back to zk.
Memory backends return shell command strings. Core asks its host to execute them and parses stdout into notes. This works through either Zellij or Montana; the required memory CLI must be available in the host’s execution environment.
Tools
The agent interacts with memory through two built-in tools:
__MEMORY_SAVE__— write a note.__MEMORY_SEARCH__— search notes.
Notes
A parsed note (MemoryNote) has:
| Field | Meaning |
|---|---|
id | identifier, usually the filename without extension |
title | title from the note’s frontmatter |
path | full path to the note file |
tags | tags associated with the note |
body | note body (may be truncated) |
score | relevance score 0.0–1.0; only set by QmdBackend |
Directory layout
Notes live under <cowboy_dir>/memory/, a zettelkasten managed by zk:
<cowboy_dir>/memory/
.zk/ # zk configuration and index
daily/ # YYYY-MM-DD.md journals
facts/ # atomic knowledge notes
decisions/ # decision records
templates/ # zk note templates
<cowboy_dir> is $HOME on a ranch install and $XDG_DATA_HOME/cowboy
(default ~/.local/share/cowboy) in lite mode.
Pre-tool retrieval
Memory search is wired in as a retrieval step before the agent runs tools,
gated by the pre_tool_retrieval config flag (default on). When enabled, the
harness searches memory based on the conversation and injects matching notes
into context so the model can use prior knowledge without an explicit search.
Retrieval does not fire inside sub-agents.
The MemoryBackend trait also exposes init_commands(), recent(limit),
search_deep(query), remember(title, content, template), journal(entry),
and an optional search_skills_deep(query) for searching skill files alongside
memory.
Sub-Agents
Sub-agents are child agent instances spawned by the harness to run a scoped task with a filtered toolset. The Zellij host opens a plugin pane; Montana starts another component instance on a thread in the same process. Both use files to communicate with the parent. Montana peers share its preopened directories and HTTP client; spawning a peer does not create another OS security boundary.
Source: crates/core/src/subagent.rs, crates/core/src/spawn.rs,
crates/core/src/subagent_config.rs.
How spawning works
The parent calls the built-in spawn_subagent tool. The harness then:
- Writes the task to
prompt.mdand the filtered tool manifest totools.jsonin a new directory under<cowboy_dir>/subagents/. - Requests a peer through the host, passing a JSON-encoded
SubagentConfigunder thesubagent_configkey. Zellij loads a visible plugin pane; Montana’s supervisor creates a component instance with its own agent ID. - The child, on its first heartbeat tick, writes a
lockfile and itspane_id, readsprompt.md+tools.json, and processes the prompt as a user message. - When the child finishes, it writes
response.mdand removeslock. - The parent detects completion on its heartbeat poll and injects the response back into its own conversation.
The shared spawn protocol uses host-executed shell commands for its spool files. WebAssembly does not itself forbid file I/O: Montana also grants direct WASI access to preopened directories. Pane visibility is a Zellij affordance; Montana has no terminal panes.
Directory layout
Each sub-agent gets a directory named <type>-<short_uuid>:
<cowboy_dir>/subagents/<type>-<id>/
prompt.md # task + metadata, written by parent
tools.json # filtered tool manifest, written by parent
lock # present while running, written by child
pane_id # host instance id (a pane id under Zellij)
response.md # final output, written by child on completion
<cowboy_dir> is $HOME on a ranch install and
$XDG_DATA_HOME/cowboy (default ~/.local/share/cowboy) in lite mode.
Sub-agent types
Three types are defined in the SubAgentType enum. Each restricts the tools
the child may call:
| Type | Allowed tools |
|---|---|
| Research | read, search, find, web-search, ls |
| Code | read, write, search, find, bash, ls, __HASHLINE_READ__, __HASHLINE_EDIT__ |
| Review | read, search, find, bash, ls |
Only Code gets write and the hashline edit tool. Review still has bash,
so its tool list does not enforce read-only filesystem access. The configured
host and tool policy determine what commands can do.
Models
SubagentConfig::cheap_defaults() selects the child’s provider and model.
Its fallback model names are:
| Provider | Default sub-agent model |
|---|---|
| OpenAI | gpt-4o-mini |
| OpenRouter | openai/gpt-4o-mini |
| Anthropic | claude-3-5-haiku-latest |
| Ollama | mistral:7b |
| Codex | gpt-5.6-luna |
| Vercel | zai/glm-5.3-flash |
An explicit subagent_model wins when that provider has a non-empty key.
Otherwise the main harness provider/model is
used when it is usable (a key, or a keyless Ollama/Codex provider), falling back
to the first keyed provider in the order OpenRouter, Anthropic, OpenAI, Vercel —
so the Ollama and Codex defaults above are only reached through the main
model.
Status
The parent tracks each child with the SubAgentStatus enum:
Starting— directory created, peer startingRunning—lockpresentCompleted—response.mdpresent,lockremovedFailed(String)—lockremoved without aresponse.md
Polling happens on the heartbeat timer.
Hashline Edit Format
Hashline is the harness’s read/edit format. Each line is tagged with a short hash of its content, so an edit can be rejected when the file has changed since the agent last read it.
Source: crates/core/src/hashline.rs. Exposed as the built-in tools
__HASHLINE_READ__ and __HASHLINE_EDIT__.
Format
Lines are formatted as {line_num}:{hash}|{content}:
1:a3|fn main() {
2:7f| println!("hello");
3:b2|}
__HASHLINE_READ__ returns a file in this form. The agent then references
lines by line:hash when editing.
Hash function
The hash is FNV-1a (offset basis 2166136261, prime 16777619) truncated to the
low byte (hash & 0xff) and formatted as two lowercase hex characters. With
256 possible values it is meant for change detection, not collision resistance.
Edit operations
__HASHLINE_EDIT__ takes an array of operations. Each references a target line
by line:hash:
replace— replace lines fromstartto optionalend(inclusive) with new contentinsert_before— insert content before the referenced lineinsert_after— insert content after the referenced linedelete— delete lines fromstartto optionalend(inclusive)
Every referenced line’s hash is validated against a fresh read of the file
before anything is applied. If any hash does not match, the whole edit is
rejected and the agent is told to re-read. Operations are applied in reverse
line order (so earlier edits don’t shift later line numbers), then the file is
rewritten in place with exactly the computed bytes — a file without a trailing
newline stays without one, and a blank last line stays. The write itself
re-checks that the file still holds the bytes that were read (cksum, in the
same shell invocation): if it changed in between, nothing is written and the
error names the stale line references. That check needs cksum on the agent’s
PATH and allowed by its sandbox policy (Bash(cksum) is in the default allow
list); if the check itself cannot run, the edit is refused with the shell’s
error rather than a retry hint. Files that are not valid UTF-8 or that contain
NUL bytes are refused, since the write cannot reproduce them byte-exact.
The parser accepts either operations or edits for the array key, and either
content or new_text for the text field.
Browser Automation
The harness can drive a real browser through Camoufox, an anti-detection Firefox fork. Agents navigate pages, read accessibility snapshots, interact with elements, run JavaScript, and take screenshots.
Source: pkgs/camoufox/, modules/camoufox.nix,
modules/options/camoufox.nix.
Components
camoufox— the Camoufox browser binary (anti-detection Firefox, daijro/camoufox).camofox-server— a Node.js REST server wrapping Camoufox (jo-inc/camofox-browser).cowboy-browser— a shell CLI (pkgs/camoufox/cowboy-browser.sh) that the agent calls; it talks to the server’s REST API.
All three are exposed as flake packages: nix build .#camoufox,
.#camofox-server, .#cowboy-browser.
Deployment
The server runs inside a shared NixOS container (cowboy-camoufox) with one
camoufox-<user> systemd service per agent. Each agent’s server listens on
basePort + (uid - 1338), where basePort defaults to 9377. Browsing happens
under a headless Xvfb display.
Enable it with:
services.cowboy.camoufox.enable = true;
This requires services.cowboy.secretsProxy.enable = true — the camoufox
container uses the proxy’s veth networking, and the module asserts this.
Options under services.cowboy.camoufox:
| Option | Default | Meaning |
|---|---|---|
enable | false | enable the camoufox container |
basePort | 9377 | base port; per-agent port derives from UID |
CLI
cowboy-browser keeps a session id in a file under $XDG_RUNTIME_DIR and
targets the server at $CAMOFOX_URL (default http://10.200.0.1:9377, the
veth host address — localhost does not reach the container). Subcommands:
cowboy-browser navigate <url> Open URL (creates a session if needed), return a snapshot
cowboy-browser snapshot [--full] Accessibility-tree snapshot of the current page
cowboy-browser click <ref> Click an element by ref
cowboy-browser type <ref> <text> Type text into an element
cowboy-browser scroll <direction> Scroll (up/down/left/right)
cowboy-browser press <key> Press a key (Enter, Escape, Tab, ...)
cowboy-browser back Navigate back
cowboy-browser eval <js> | eval - Run JS in the page ('-' reads from stdin)
cowboy-browser screenshot [file] Save a screenshot (default /tmp/cowboy-browser-screenshot.png)
cowboy-browser close Close the browser session
Elements are referenced as @e1, @e2, … from the snapshot output (the @
is optional).
Example
cowboy-browser navigate https://example.com
# snapshot lists interactive elements with @eN refs
cowboy-browser type @e3 "search query"
cowboy-browser click @e7
cowboy-browser screenshot ./result.png
The agent invokes these through its bash tool. cowboy-browser unsets the
HTTP proxy variables for its own requests, since the camofox server is local to
the veth and does not need credential injection.
API Integration
Provider code formats requests and parses responses in core. The host carries them out asynchronously: the Zellij adapter uses its web-request API, while Montana implements the component’s HTTP import with a native client and sends the completion back as a host event. Provider logic does not depend on a terminal or Zellij panes. Proxy routing, credentials and network policy still depend on the deployment.
Source: crates/core/src/provider/.
Providers
Six providers implement the LlmProvider trait, selected by ProviderType:
ClaudeProvider(Anthropic Messages API)OpenAIProvider(Chat Completions / Responses API)CodexProvider(ChatGPT subscription Responses API with SSE)OpenRouterProviderVercelProvider(Vercel AI Gateway, Chat Completions)OllamaProvider(local, keyless)
The LlmProvider trait
The trait does not perform transport I/O. It only serializes requests and deserializes responses:
#![allow(unused)]
fn main() {
pub trait LlmProvider {
fn name(&self) -> &str;
fn format_request(
&self,
messages: &[Message],
tools: &[Tool],
system: &str,
) -> (String, BTreeMap<String, String>, Vec<u8>);
fn parse_response(&self, body: &[u8]) -> Result<LlmResponse, ProviderError>;
fn parse_error(&self, status: u16, body: &[u8]) -> ProviderError;
fn set_model(&mut self, model: &str);
fn api_key(&self) -> &str;
}
}
The request flow is:
format_request()produces(url, headers, body).- Core requests HTTP through its host interface.
- The host delivers the correlated HTTP result: a Zellij event or, through Montana, a WIT host event mapped back into core.
parse_response()(on HTTP 200) orparse_error()(otherwise) interprets it.
set_model() supports switching the model at runtime (the /model command).
Credentials
In the managed credential-proxy deployment
(services.cowboy.secretsProxy.enable), the agent sends placeholder keys and
the proxy injects the real credentials on the wire based on per-domain mappings.
Codex is the same pattern: its rotating ChatGPT subscription token is injected
outside the agent. Local and ordinary Docker launches hold credentials.
See Security Model. For keyless
providers (Ollama) api_key() is empty.
Provider defaults
ClaudeProvider targets /v1/messages. Its defaults: model
claude-sonnet-4-20250514, max_tokens 8192, extended thinking enabled with a
token budget. OpenAIProvider defaults to gpt-4o and auto-selects the
Responses API for reasoning models (o-series, gpt-5, codex) to capture reasoning
summaries, which are mapped onto thinking content blocks.
Message model
Messages carry a string role and a vector of ContentBlocks, matching the
Claude API’s native content-block model. A block’s block_type is one of
text, tool_use, tool_result, or thinking.
Errors and retries
ProviderError includes ParseError, ApiError, NetworkError,
RateLimited, Timeout, InvalidRequest, AuthenticationError, and
Overloaded. is_retryable() marks rate limits (429), network errors, server
errors (5xx), and timeouts as retryable; auth and other 4xx errors are not.
Automatic retry is wired in: a RetryState with exponential backoff drives a
bounded retry of failed LLM calls from the ranch handlers
(crates/core/src/handlers.rs, schedule_llm_retry) — a retryable error
schedules a backed-off retry until max_attempts is exhausted, after which the
failure surfaces as an error message.
Bridges
Bridge services (Discord, email) do not use the provider layer. They reach the agent through the pub/sub message system: inbound messages are polled and processed as user input, and replies flow back out the same way. See Approvals & Outbox.
Plugin Architecture
title: Plugin Architecture tags: [architecture, plugins, extensibility] created: 2026-04-24 updated: 2026-09-06
Plugin Architecture
Cowboy’s NixOS module system is the extension surface. An external module
(a “plugin”) adds capability to agents by registering into cowboy’s
option tree rather than forking the core. The in-tree
Home Assistant skill is the
worked example of per-agent content; the examples below use a hypothetical
myplugin that runs a service on the host and teaches an agent to drive it.
Everything below is grounded in the implementation under modules/.
The three layers
A plugin separates into layers so that only the top one couples to cowboy:
| Layer | What it is | Couples to cowboy? |
|---|---|---|
| Infrastructure | systemd services / scripts providing the raw capability (game server, DB, hardware) | No — runs standalone |
| Packages | derivations giving the programmatic interface (Python env, CLI) — built via callPackage | No — self-contained |
| Harness integration | conditional blocks registering skills/tools/bridges and pushing packages into agent homes | Yes — guarded by hasCowboy |
In myplugin this maps to: services.nix (infrastructure), default.nix
(packages, built once and imported by both of the others), and the config
block in module.nix guarded by hasCowboy (integration). Disable cowboy
and the service still runs.
The plugin contract
1. Own your namespace
Declare options under your own top-level namespace, never inside
services.cowboy.*:
# Good
options.services.myplugin = { enable = ...; dataDir = ...; };
# Bad — destabilizes cowboy's option tree and breaks standalone use
options.services.cowboy.myplugin = { ... };
2. Read cowboy state via cowboyLib
Cowboy injects a cowboyLib module argument
(modules/lib/default.nix). This is the supported module interface today, not a
versioned binary ABI or a promise of compatibility across future revisions.
Pin the Cowboy input and evaluate extensions when updating it. Members:
cowboyLib.enabledAgents # attrset of enabled agents (name -> agentConfig)
cowboyLib.agentNames # [ "agent" ... ] — for building `agents = [ ... ]` selectors
cowboyLib.forAgents (acfg: {…}) # map an HM config fn over every enabled agent, keyed by user
cowboyLib.forAgentsWhere pred f # …only agents matching `pred name acfg` (generic haAgents)
cowboyLib.singleAgent "context" # the one enabled agent, or a lazy throw under multi
cowboyLib.broker # human operator user (may be null)
cowboyLib.serviceUser # cowboy's service-account prefix / shared bridge group
cowboyLib.userFor / homeFor name # agent user / home dir by name
cowboyLib.caBundleFor name # the trust store the agent's TLS clients must use ($SSL_CERT_FILE)
cowboyLib.inNamespace # is agent traffic proxied?
Accept it with a fallback so the plugin still evaluates when cowboy is absent:
{ config, lib, pkgs,
cowboyLib ? { enabledAgents = {}; forAgents = _: {}; },
... }:
let hasCowboy = cowboyLib.enabledAgents != {}; in
Per-agent targeting:
forAgentsand theskills/toolsregistries fan out to all enabled agents by default. To scope a skill/tool to specific agents, set itsagents = [ "name" … ](empty = all). For per-agent content (prompts that differ by agent), useforAgentsWheredirectly — see the Home Assistant pattern.
3. Guard harness integration
Keep infrastructure unconditional; gate cowboy registration on hasCowboy:
config = lib.mkIf cfg.enable {
# Infrastructure — always
systemd.services."myplugin-setup" = { ... };
# Integration — only with cowboy
services.cowboy.skills = lib.mkIf hasCowboy { ... };
home-manager.users = lib.mkIf hasCowboy (
cowboyLib.forAgents (acfg: { home.packages = [ pluginCli ]; })
);
};
4. Don’t block cowboy.target
Plugin services join the agent lifecycle with a soft pull-in:
- Use
wantedBy = [ "cowboy.target" ]for startup with the target. Add stop and restart coupling only when the plugin needs that lifecycle. - Bound restart loops with
StartLimitBurstandStartLimitIntervalSecwhen usingRestart = "on-failure". - Use
requires/after/partOfamong the plugin’s own services for internal ordering.
5. Keep packages self-contained
callPackage your own dependencies, and build each one once — a
default.nix that both module.nix and services.nix import, rather than
two copies of the same derivation kept in sync by hand. Don’t assume cowboy
provides any particular package on PATH (the base set is just bash coreutils ripgrep fd jq gh nix git curl, plus camoufox’s browser — see
modules/tools.nix).
Extension points
Skills — knowledge + prompt + packages
A skill is a markdown prompt the agent can load, optionally with packages
and extra tools. Registered into the global services.cowboy.skills
attrset (modules/skills/default.nix):
services.cowboy.skills.myplugin-guide = {
description = "Drive myplugin from the CLI";
prompt = ./skills/myplugin-guide.md; # or promptText = "...inline...";
requires = [ pluginCli ]; # installed for every targeted agent
additionalTools = [ "bash" "read" "write" ];
tags = [ "myplugin" "control" ];
autoLoad = false; # true = loaded at agent startup
agents = [ ]; # [] = all enabled agents; or [ "pilot" ]
};
For each agent the skill targets, the module writes
~/.config/cowboy/skills/<name>.md (the harness discovers skills by
reading this directory) plus a per-agent skills.json, and installs the
skill’s requires packages onto that agent’s PATH — regardless of
autoLoad, so an on-demand skill’s binaries are present when the agent
loads it. The agents selector filters which agents receive the skill.
Tools — schema’d executable capabilities
A tool is a command template the harness can call with structured args
(modules/tools.nix). The default set (bash, web-search,
spawn_subagent) merges with anything a plugin adds:
services.cowboy.tools.myplugin-eval = {
package = pluginCli;
command = "myplugin run {{script}}"; # placeholders are {{double-brace}}, NOT {single}
description = "Run a myplugin control script";
sandbox = "standard"; # "none" | "standard" | "strict"
timeout = 60;
agents = [ ]; # [] = all enabled agents (defaults always reach every agent)
schema = {
type = "object";
properties.script = { type = "string"; description = "Path to script"; };
required = [ "script" ];
};
};
Tools are baked per-agent into the harness WASM and written to a per-agent
~/.config/cowboy/tools.json; the agents selector scopes a tool to
specific agents (the default tools — bash, web-search,
spawn_subagent — leave it empty, so they reach everyone). A plugin can
also expose a capability through a skill alone (prompt + its binary on
PATH), which is the lighter path when the agent only needs a CLI.
Plugin registry — be discoverable
Register the plugin in the informational registry so the agent (via its
system prompt) and the operator (cowboy plugins / /etc/cowboy/plugins.json)
can enumerate what’s installed:
services.cowboy.plugins.myplugin = {
description = "Host service control";
version = "1";
extensionPoints = [ "skills" "units" ];
units = [ "myplugin-setup.service" "myplugin-daemon.service" ]; # the real units it owns
agents = [ ]; # [] = surfaced to all agents
};
This is purely descriptive — it wires up no capability, it just makes the
plugin enumerable. The per-agent tools.json system_prompt_suffix
carries the list to the agent with no harness changes.
Bridges — bidirectional message channels
A bridge connects an external platform to the agent’s pubsub
(modules/bridges.nix, options in modules/options/bridges.nix). Each
declaration auto-generates a pubsub source plus three systemd units
(cowboy-<name>-{ping,ingest,outbox}) under cowboy-bridges.target:
services.cowboy.bridges.discord = {
pkg = myBridgePkg; # provides bin/discord-{ping,ingest,outbox}
env = "/run/agenix/discord-env"; # EnvironmentFile
icon = "💬";
approval = {
required = true;
notify = "discord"; # which outbox sends the approval prompt
notify_channel = "1489..."; # channel id within that outbox
timeout = 3600; # auto-reject after N seconds
};
};
Bridges default to dedicated service accounts. Keep their identities separate
from agents and the operator when overriding user. Configure outbound
approval on the bridge and inbound routing with routes / defaultAgent.
The Security Model explains the separate Redis authorities.
The systemd target
A unit can be started with the agent target:
systemd.services.my-capability.wantedBy = [ "cowboy.target" ];
Per-agent skills: the Home Assistant pattern
The global skills registry distributes to all agents by default and supports
an agents selector. For content that differs per agent, the in-tree HA skill
(modules/skills/home-assistant.nix) shows the escape hatch: it
bypasses the registry and writes the prompt directly into the homes of
agents that opted in via a per-agent option (agentOpts.skills.homeAssistant):
let haAgents = lib.filterAttrs (_: a: a.skills.homeAssistant.enable) enabledAgents; in
home-manager.users = lib.mapAttrs' (_: acfg:
lib.nameValuePair acfg.user {
home.file.".config/cowboy/skills/home-assistant.md".source = mkHaSkillPrompt acfg;
home.packages = [ pkgs.curl pkgs.jq ];
}
) haAgents;
The agents selector handles per-agent selection directly in the
registry, so a plain skill that only needs targeting does not need this
bypass. HA uses it because its prompt is per-agent content — the
endpoint, token, and managed units differ per agent — which a
selection-only list can’t express. Use cowboyLib.forAgentsWhere for that
case; HA is the worked example.
Worked example: myplugin end to end
myplugin/
├── module.nix # options + integration layer (skills, agent packages, registry entry)
├── services.nix # infrastructure: myplugin-setup / myplugin-daemon units
├── default.nix # packages: the CLI and its client library (callPackage), built once
└── skills/*.md # skill prompts
Consumer wiring (machines/<host>.nix):
imports = [
inputs.cowboy.nixosModules.default
../modules/myplugin/module.nix
];
services.myplugin = {
enable = true;
dataDir = "/home/<user>/myplugin";
user = "<user>";
};
What lights up because cowboy is present:
- The plugin’s skills, written into the agent’s
~/.config/cowboy/skills/. pluginClion the agent’s PATH.- The agent’s polkit-managed unit list (
managedServices.units) so it can cycle the stack itself.
Managed units must exist. Every name in
managedServices.unitsmust be a unit the configuration declares; a polkit rule over a unit that does not exist advertises a control that silently does nothing.modules/services.nixchecks this at eval time and warns rather than asserts (services.cowboy.managedServices.validateUnits, defaulttrue), because a host’s units may be legitimately conditional — a feature flag, an optional submodule — and a hard assertion would force every consumer to mirror that logic into its unit list. The registry’splugins.<name>.unitslist is informational and is not validated.
Checklist for a new plugin
- Options under your own namespace, with an
enableflag. -
cowboyLibaccepted with a{ enabledAgents = {}; forAgents = _: {}; }fallback. - All cowboy registration guarded by
hasCowboy. - Packages via
callPackage, built once indefault.nix; not assumed present. - Plugin services declare startup dependencies and bounded restart behavior.
- A skill’s binaries go in its
requires(installed for every targeted agent, on-demand or not). - Scope to specific agents with
agents = [ … ]; for per-agent content, useforAgentsWhere. - Register in
services.cowboy.plugins.<name>so the plugin is discoverable. - Every unit in
managedServices.unitsactually exists (the validator warns otherwise);plugins.<name>.unitsis informational.
Different contracts
The Nix module interface configures deployments. The component interface is
separately versioned as cowboy:agent@0.2.0; its WIT source defines commands,
effects and carriers. The JSON inside a snapshot remains an opaque display
model, not a versioned presentation schema. Native control and descriptor
records have their own CDDL definitions.
The A2A adapter exposes a limited profile over the root agent. It does not provide general A2A task persistence, peer addressing or every protocol method. See the Component ABI and A2A profile.
Design Overview
Cowboy runs persistent AI agents on infrastructure you control, using Nix to configure their tools, credentials, message access, and approved actions.
The strongest documented deployment is managed NixOS on Linux/x86-64. Local and ordinary Docker launches have different credential and process boundaries. The OCI runtime is a specialized optional route for Cowboy’s own payload. See Deployment Paths for that choice before selecting a host or changing its controls.
Core and the two hosts
The core owns the agent loop, provider request formatting, tools, session state, memory and message polling. Two hosts run it:
| Host | Core integration | Interaction |
|---|---|---|
| Zellij plugin | Links core into a wasm32-wasip1 plugin | Terminal input and ANSI display |
| Montana | Loads a wasm32-wasip2 component through cowboy:agent@0.2.0 | Commands in and JSON frames out over the native control socket |
Zellij implements asynchronous commands, HTTP requests, timers and peer spawning through its plugin API. Montana implements those operations natively and delivers completions to the component. Both use the same semantic intents and provider logic; terminal navigation and pane controls belong to Zellij. Montana runs peers as component instances on separate threads in the same process, sharing its filesystem view and HTTP client.
Two paths to the outside world
Custom asynchronous effects cross the core’s host interface: execution, HTTP, timers and peer lifecycle requests. The component adapter maps these to WIT imports. An effect request does not grant permission to perform it.
Direct filesystem access is a separate path. Montana gives the component WASI preopens for its resolved workspace, home, config, data and temporary directories. Config loading, plan persistence and other direct file operations use those preopens without an execution effect. WASI clocks and randomness are also available; WASI sockets are not enabled.
Montana maps each preopened host directory to the same guest path. This matters because execution effects run native processes: a path computed in the guest must name the same file when passed to a shell command.
To reason about access, inspect both the host’s effect handling and its WASI grants, then the enclosing process identity, mounts and network policy. Sheepdog’s tool-command policy alone does not describe direct guest file access. See Security Model.
Launch branches
cowboy-core
├── Zellij plugin: interactive terminal
└── WASI component → Montana
├── managed NixOS: systemd agent service
├── local headless launch
├── ordinary Docker: image supervisor
└── optional Cowboy OCI runtime: embedded Montana
The managed daemon runs Montana directly with a per-agent component, under the shared systemd hardening envelope, agent user, configured network namespace, credential proxy and Sheepdog policy. It does not launch the agent through the OCI runtime. Optional socket forwarders use a helper from that runtime package; that does not turn the daemon into an OCI container.
Local and ordinary Docker launches hold provider credentials. The managed proxy deployment keeps them outside the agent. The OCI route has its own mount, seccomp and egress controls; it is not a general container runtime. Tool behavior can differ under these policies. The terms bare, user and container classify a design progression; they are not three switches offered by the runtime binary.
Messages and configuration
Bridge services connect external platforms to Redis streams. Core polls its configured sources and sends replies through their outboxes. Bridge approval can hold an outbound message for a human decision. Recovery does not guarantee completion or preserve reply routing across a restart; see Inbound Message Reliability.
Both hosts pass a flat string configuration map to core. The CLI, persistent JSON, generated NixOS configuration and secret discovery are different inputs; see Configuration. Skills, tools and bridges extend the agent through the module interface.
Implementation map
| Area | Location |
|---|---|
| Agent state and tools | crates/core/ |
| Terminal host | crates/harness/ |
| WIT adapter | crates/component/ |
| Headless host and WASI grants | crates/montana/ |
| Specialized OCI runtime | crates/runtime/ |
| Native protocol types | crates/contracts/ |
| CLI and launch selection | cowboy/ |
| Managed services and policy | modules/ |
| Credential injection | proxy/ |
| Message bridge framework | pkgs/bridge/ |
The Component ABI describes WIT and the native control protocol. The embedding guide explains how to add a host.
Security Model
The strongest documented deployment is managed NixOS on Linux/x86-64. Its agent daemon runs Montana directly under systemd with a per-agent component, an agent user, the shared hardening envelope, a network namespace, the credential proxy and Sheepdog. It does not use the OCI runtime to launch the agent. Local, ordinary Docker and optional OCI launches have different boundaries; see Deployment Paths.
This page explains the controls. The threat model states their conditions and limitations. In particular, Cowboy does not guarantee confidentiality of anything the agent can read. A prompt-injected worker with permitted outbound access can disclose readable data. The agent can also damage its own writable workspace, memory and state.
What is enforced by default
The Nix modules enable the credential proxy by default, and enable Redis ACLs and Sheepdog by default on Linux. These are deployment settings, not properties of an arbitrary core or component build. Managed daemons require the proxy; turning it off is not a supported way to run that daemon with real keys.
Process and filesystem isolation
Each managed daemon runs as its configured agent user. The systemd envelope makes the system read-only, hides other homes, binds the agent’s home back in, and applies private temporary storage, device restrictions, an empty capability bounding set and no-new-privileges. Configured mounts add access to that view. The home, runtime and explicitly writable paths remain mutable.
Inside Montana, direct guest file access uses WASI preopens. These map the workspace, home, config, data and temporary directories at their host paths, so guest reads and native execution agree about filenames. Preopens grant file access separately from custom execution effects; a review of tool policy alone is not a review of all filesystem access.
Montana sub-agents share the host process and its preopens. A filtered child tool list is not a separate OS identity or filesystem boundary.
Network isolation
The managed Linux module creates one shared namespace, cowboy-ns, with a
veth pair to the host. The daemons and agent user services join it. It is not
one namespace per agent.
Outbound TCP outside the configured subnet is redirected to the proxy. The egress filter allows loopback, the subnet, established replies and DNS to the host resolver, then drops other traffic. Operator-configured SSH destinations and passthrough ports grant direct egress exceptions.
The host Nix daemon is another network authority: its builders run outside the agent namespace. The default denial of direct Nix daemon access is a Sheepdog connect rule. Disabling Sheepdog removes that enforcement. Rebuild requests can instead go through the broker’s configured approval workflow.
Darwin uses a UID-scoped packet-filter redirect and has neither the Linux namespace boundary nor Sheepdog or the Linux Redis ACL implementation.
Credential-injecting proxy
Managed daemons require the proxy and use placeholder provider keys. The proxy reads the real secret from its configured file and injects it on a matching TLS destination route. Exact host routes take priority over wildcard routes; more specific wildcard suffixes take priority over broader ones. A mismatched Host header is rejected before credentials are attached. Upstream certificate verification is enabled by default; the explicit insecure option disables it.
These claims apply to the proxy deployment. Local and ordinary Docker agents hold credentials. Proxy configuration does not remove credentials supplied through some other input.
With egress control enabled, GET, HEAD, OPTIONS and TRACE pass the method gate. Every other method needs an allowed destination or receives HTTP 403. These are method-and-destination restrictions, not a guarantee that external state cannot change. There is no content inspection that prevents readable data from leaving through an allowed request.
Sheepdog policy
On Linux/x86-64, the managed build bakes a per-agent Sheepdog policy into the component’s tool execution path. Sheepdog uses seccomp notification to mediate selected file, execution and connection syscalls. Its policy includes read, edit, create, delete, execution and connection rules. Runtime permission grants can take precedence over baked rules when enabled.
This is conditional on the sandbox being enabled and present in the build. Portable lite builds do not enforce this policy. The optional OCI route uses its own bundle restrictions; it does not inherit all managed-service controls. The policy authority and limitations are specified in Sheepdog policy.
A command can exit successfully even when part of it was denied. Core preserves recognized Sheepdog denial messages in the tool result so the agent can see that a refused read or connection was not an empty successful result.
Per-agent message isolation
With the Linux Redis ACL configuration enabled, each agent has its own inbox and outbox streams and an identity restricted to its name prefix. It cannot read another agent’s streams or edit approval records.
The agent reaches Redis through its own Unix socket at
/run/cowboy/<agent>.sock, owned by its user with mode 0600. The host auth proxy
selects the Redis identity from the accepting socket and injects the password.
The agent holds no Redis password. Per-agent source configuration names these
streams and the socket.
Bridges route inbound messages to an agent’s inbox and consume each agent’s outbox. The outbox stream prefix determines the requester; a message’s claimed user does not. Bridge bookkeeping lives outside agent-writable prefixes. Each bridge has its own service user and Redis identity, scoped to its streams and required approval roles. A bridge that creates or resolves approvals still has authority over shared approval records; those roles must be trusted.
This separates agent mail while provider credentials remain shared through the host’s proxy. It is not a tenant boundary between mutually untrusted operators.
Approvals and operational authority
Bridge approvals hold configured outbound messages for a human decision and reject them on timeout. Declared Git effects bind approval to a host-inspected commit and base; apply rechecks policy before pushing. The rebuild bridge has privileged activation authority through its configured wrappers. Treat bridge identities and approval resolvers as part of the trusted host.
See Approvals & Outbox and Gated Effects. After a restart, inspect outcomes before retrying: message recovery can repeat work without restoring the reply’s source binding.
Agent users do not receive journal access by default. Enabling
services.cowboy.agents.<name>.allowJournal grants access to host logs as well
as the corresponding log-path policy. Logs can contain data from other trust
domains. The proxy omits query strings from its own request logs; that does not
sanitize output from other services.
System Patterns
Keep credentials behind a service
In the credential-proxy deployment, provider secrets live outside the agent. The agent sends placeholders; the proxy selects the destination route and injects the credential. Local and ordinary Docker launches deliberately hold credentials, so this is a property of the proxy deployment, not of core.
The proxy restricts HTTP by method and destination. With egress control enabled, GET, HEAD, OPTIONS and TRACE pass the method gate; every other method requires an allowed destination. This does not prove that an allowed request leaves external state unchanged.
Cowboy does not guarantee confidentiality of data the agent can read. A prompt-injected worker with permitted outbound access can disclose that data. Keeping provider keys outside the worker does not make its workspace private. See Security Model.
Separate proposal from authority
An agent can prepare a commit or request a rebuild. The bridge and configured host service perform the approved action with their own identity. For declared Git effects, the approval binds the commit and base that the host inspected; the apply step checks them again. See Gated Effects.
Extend configuration before core
Tools, skills, packages and bridges can be registered through Nix modules. Per-agent selectors determine who receives them; the current module interface provides accessors for enabled agents and their homes. Adding such an extension does not require changing the agent state machine. See Plugin Architecture.
Separate shared behavior from host affordances
The two hosts share agent state and semantic commands. Zellij owns terminal navigation; Montana owns the native control socket and component instances. A host must supply both effect handling and the filesystem access its guest uses. Host independence does not imply identical security boundaries.
Component ABI
The agent core (cowboy-core) shares its state machine across hosts. Its
custom asynchronous interface is effects out (a Host trait) and
inputs in — split into
semantic Intents (the shared vocabulary) and raw HarnessEvents (the Zellij
transport plus effect completions). This chapter is its ABI: the WIT world
cowboy:agent@0.2.0 at
specs/wit/cowboy-agent.wit
— the single WIT source every bindgen consumer reads — which lets a runtime
embed core without linking Rust in-process.
The world is small enough to state in a sentence. It exports one interface,
agent, with five functions — initialize, handle-command,
handle-host-event, snapshot, describe — and imports one, host, for
the effects a component requests: exec, http, set-timer, close-self,
spawn-peer, and own-agent-id.
Two host paths
Core has two adapters, and only one goes through WIT:
| Host | Crate | Path to core | Input | Output |
|---|---|---|---|---|
| Zellij plugin | cowboy-harness (wasm32-wasip1) | links cowboy-core natively, calls the Rust trait | keys/mouse/paste → navigation stays local, semantic keys → Intent | ANSI to stdout (Zellij captures) |
Headless embedder montana | cowboy-component (wasm32-wasip2) | WIT cowboy-agent world | command + host-event | JSON display model via snapshot |
The Zellij adapter deliberately bypasses the component ABI. It needs keystroke-level fidelity to drive the terminal UI (prompt-line editing, navigation, expand/collapse), and it can link Rust, so paying the component-boundary tax buys it nothing. The WIT world exists for hosts that cannot link Rust — and those hosts don’t remote a TUI, they render the JSON display model and build their own input affordances.
The TUI lives in the adapter.
cowboy-coreis a headless agent model: it owns theDisplayItemdata andAgentHarness::frame_json, while the renderers, syntax highlighting (syntect), modal input, and view state (ViewState) live incrates/harness/src/tui— wrapped around the core asTui { agent, view }, which derefs toAgentHarness. The generic ANSI atoms (colors, symbols, text, spinner, key types) live incowboy-ui. Core depends on neithercowboy-uinorsyntect, so the WASI component sheds both. Every view↔core coupling crosses theHostseam as a defaulted hook (below), never a field poke.
The named-intent seam
Raw keystrokes never cross the ABI. Core’s input surface splits in two:
- Navigation / view — scroll, expand/collapse, search, prompt-line editing. Client-local: the Zellij adapter owns it (it holds the TUI projection); a JSON client navigates its own copy of the frame. Never crosses WIT.
- Semantic intents — the handful of actions that change agent/session state:
submit/interrupt/approval/debug/quit. The type isIntent, owned bycowboy-contractsand re-exported bycowboy_core::intentso the native adapters and the component adapter cannot drift.
Both adapters converge on one vocabulary via AgentHarness::handle_intent —
no bifurcation. The Zellij adapter’s semantic key arms (in
crates/harness/src/tui/input: prompt.rs Enter → Submit, navigate.rs
Esc → Interrupt and d → Debug, mod.rs sentry y/o/A/n →
Approval) call the same method the WIT component maps command’s cases onto.
Intent derives serde, so it is also the montana driver’s inbound wire shape
({"submit":"hi"}, "interrupt", {"approval":"once"}, "debug", "quit",
specified by specs/control.cddl)
— one vocabulary end to end, not a second protocol.
Commands vs. host events
The world’s inbound surface is two types, delivered through two exports, because they have different producers and different trust semantics — a transport command must never be confusable with an effect completion:
command— semantic operations an operator or protocol adapter applies to the instance:submit(string),interrupt,approval(approval),shutdown. Delivered byhandle-command.host-event— completions and notifications the host produces:timer,command-result,web-result,permission-result. Delivered byhandle-host-event.
Both return bool: true when observable state changed and a fresh snapshot
is worth pulling.
debug is deliberately not a WIT command — it opens a local log viewer, so
it is a host diagnostic rather than an agent operation, and it stays in the
native Intent vocabulary only. Intent::Quit is what command.shutdown maps
onto guest-side.
Neither type carries key, mouse, or paste. Core keeps the full
Key/KeyEvent/Pasted + HarnessEvent types for the Zellij transport.
One component instance has one foreground turn. The world exposes no task
handle, no detach/background operation, and no reattachment; submitting while
work is active queues the input, and interrupt asks the guest to abandon its
turn without promising cancellation of already-issued host effects. Peers
spawned via host.spawn-peer are separate instances, not background tasks of
the spawner.
Rust-only host verbs
Some seams stay in the Rust Host trait and never enter the world at all,
because they are meaningless off a terminal.
The pane verb is one: pane visibility is client-local view state a windowed
host acts on, so montana/remote never needs it. pane_op lives in the Rust
trait (the native Zellij adapter implements it); the WASI component no-ops it.
Defaulted view hooks
Four more Rust-only Host methods carry the view couplings — all defaulted,
so headless hosts inherit them for free (they carry no meaning off a terminal),
and only ZellijHost overrides them:
reveal_latest()— “a new item was appended; snap the view to it.” Core calls it wherever the view must follow appended output;ZellijHostrecords a bit itsTuidrains after the event (honoring the user’s mode).open_debug(path)—ZellijHostopens the debug log in a floating pane; headless no-ops. Keeps thezellij action new-pane … less +Fexec out of core.has_draft_input() -> bool— core’s idle judge/autoreprompt guard; the draft lives in the adapter’s view state, so the default isfalse.editor_active() -> bool— core’s heartbeat cadence; the$EDITORround-trip is a Zellij-only affordance, so the default isfalse.
ZellijHost and its Tui share an Rc<HostSignals> cell block: core→view for
reveal_latest, view→core for has_draft_input/editor_active. Like pane_op,
none of these are in the WIT world — they are Rust-trait seams a windowed host
implements and everyone else ignores.
The opaque-correlation rule
Every effect that expects a reply carries a correlation — an opaque
list<tuple<string, string>> created by the guest (core holds it as Ctx, an
ordered map, and converts at the boundary). The host echoes it back verbatim
on the matching host-event and never reads or writes it. The request “kind”
lives inside it as a string, meaningful only to core (see dispatch.rs).
A correlation identifies an effect callback, not a foreground or background agent task — it is not a handle a host can hold, poll, or reattach to.
The payoff: new command/web kinds are new correlation string values, not new ABI surface. The world never changes when core grows a new internal effect kind. A host that treats the correlation as bytes-in/bytes-out is forward-compatible by construction.
Effects are requests, not grants
An exec or http call across the host import is a request. Nothing about
crossing the boundary authorizes it: the host evaluates it under the active
backend and capability policy (sheepdog, the credential proxy, the OCI
sandbox — see Security) and may refuse. The same is true of
initialize’s config map: its entries are configuration, not capability
grants. own-agent-id is the one purely informational import — the instance’s
own id, which core’s heartbeat uses to address itself.
Additive timer semantics
host.set-timer(secs) is additive and one-shot: each call arms an
independent timer, and there is no cancellation. This mirrors Zellij’s
set_timeout, which has no cancel API. Core’s heartbeat (heartbeat.rs)
compensates with an armed_count that culls the resulting fan-out down to a
single live chain.
A host must not “improve” this to one cancelable timer. Core’s re-arm logic
assumes additive delivery; a cancel-on-rearm host would stall or double-fire the
heartbeat. A headless embedder implements this as a plain deadline min-heap:
push on set-timer, pop-and-deliver when the earliest deadline passes.
The JSON frame contract
snapshot() returns one line of literal JSON:
{"status": "WaitingForInput", "items": [ /* DisplayItem… */ ]}
status is AgentStatus; items is the DisplayItem list core also renders
to ANSI. ANSI highlight caches (highlighted_lines, highlighted_input) are
#[serde(skip)]’d — they are render artifacts, not state.
Deliberately unstabilized. WIT specifies only the carrier — a string holding one JSON object. The shape inside tracks core’s display model and will change as that model settles; a typed, versioned presentation contract is deferred work, so no host may treat this serialization as a stable interface by accident.
montana’s ndjson stream wraps these frames per agent as{"seq", "agent", "frame"}— that envelope is specified, by thecontrol-framerule inspecs/control.cddl.
The describe contract
describe() returns the protocol-neutral agent descriptor, also as one JSON
object in a string: name, description, version, provider, model, capabilities,
and the skill list (one entry per tool, carrying its JSON Schema input schema
verbatim). Unlike the snapshot, this record is specified — by
specs/descriptor.cddl,
with Descriptor in cowboy-contracts as the Rust shape.
It exists so a host can advertise the agent without knowing anything about
Cowboy’s internals: crates/a2a maps this one export onto an A2A agent card and
its skill list, and montana serves that mapping. Protocol adapters are bindings
onto describe + handle-command, not alternative agent APIs.
/connect — a harness as a client of another harness
/connect <host:port>#<secret-path> (or /connect <montana-uds-path>) turns a
Zellij harness into a thin client of a remote agent, built entirely on the two
contracts above: inbound, montana’s ndjson envelope rehydrates into
Vec<DisplayItem> (pinned by display_item_round_trips_through_json);
outbound, every user action is one Intent wire line (pinned by
intent_wire_shapes). The fragment names a bearer-token file — the token
is the first line of the connection, exactly the cowboy-runtime __forward
handshake; the token itself never enters plugin config or argv.
Mechanically it is drop-and-connect: the plugin writes a request JSON to
connect_request_path and quits; the Python launcher’s relaunch loop
(cli._run) starts a fresh Zellij whose plugin sees connect_target at
load() and comes up in client mode (crates/core/src/client.rs) — no
providers, no session, no API keys. Because a WASM plugin cannot hold a
socket, a cowboy _connect bridge process owns the connection and streams
each frame line in via zellij pipe --name connect-frame; sends go out as
one-shot cowboy _connect --send connections. /disconnect is the same
sentinel dance back to a local session. Standalone entry:
cowboy connect TARGET#SECRET-PATH.
Caveats: in a full NixOS (netns) deployment, outbound TCP is DNAT’d to the
secrets proxy — a /connect target port must be in
secretsProxy.passthroughPorts (UDS targets are unaffected; sheepdog always
allows AF_UNIX), and the in-session handoff loop exists only in lite launches
(full mode: run cowboy connect from a host shell — a client session needs no
agent config anyway). A read-only __forward --ro port works as an observer:
frames flow, intents are dropped server-side.
Known Zellij leaks in core
A few Zellij-isms survive in cowboy-core. None break a headless host; they
degrade gracefully and are tracked here for eventual promotion to proper host
verbs. (The two obvious candidates are already out: the $EDITOR round-trip is
Tui-local, and the debug pane is the open_debug host hook.)
- Subagent wasm path (
spawn.rs) — the fallback peer path assumes a Zellij plugin layout (data_dir()/zellij/plugins/cowboy-harness.wasm). A headless host overrides it via config /AGENT_HARNESS_WASM. - Model-facing prose (
config/loaders.rs) — thespawn_subagenttool description says “in a separate Zellij pane”. Cosmetic; visible to the model, not load-bearing.
Filesystem access alongside WIT
Core also uses direct filesystem operations for configuration, plan state and metadata. Montana grants these through WASI preopens, mapping host directories to identical guest paths so native execution effects see the same files. These grants are separate from the custom WIT imports. An embedder must review both the preopens and effect policy; see Design Overview.
Hacking
Cowboy’s core is a headless state machine shared by the Zellij plugin and Montana. Hosts supply asynchronous effects and the filesystem access the guest uses. The optional OCI runtime embeds Montana. This guide explains where to extend that implementation.
| You want to | Start from |
|---|---|
| Run the agent under your own UI, transport or runtime | The WIT world and templates/host |
| Ship the agent to a machine without linking Rust | packages.agent-component |
| See a complete host with threads, sockets and peers | crates/montana/ |
| Add an effect the agent can perform | The Host trait in crates/core/src/host.rs |
The contract: the WIT world
specs/wit/cowboy-agent.wit defines cowboy:agent@0.2.0, the one interface
every non-Zellij host programs against. It exports agent — initialize,
handle-command, handle-host-event, snapshot, describe — and imports
host — exec, http, set-timer, close-self, spawn-peer,
own-agent-id. Commands and host events are separate exports so a transport
message can never be mistaken for an effect completion; effect results carry
an opaque correlation the host echoes back unchanged; and an effect request
is not an authorization, so the host stays responsible for policy. The
Component ABI chapter walks through each of those
rules; the file itself is short enough to read first.
WIT defines the custom interface; the guest also needs WASI filesystem grants.
The JSON inside snapshot is an opaque display model, not a stable schema.
The config map given to initialize is the flat string map core parses;
crates/contracts/src/keys.rs inventories recognized keys. A new host-provided
capability needs an import and implementations in the affected adapters.
The artifact: packages.agent-component
Build from source; see Installation for
release availability. nix build .#agent-component produces share/cowboy/cowboy-agent.wasm, the
agent core compiled to wasm32-wasip2 and wrapped by crates/component/ so
it exports exactly the world above. The wrapper, crates/component/src/lib.rs,
is small: it maps initialize onto AgentHarness::on_load, handle-command
onto AgentHarness::handle_intent, handle-host-event onto
AgentHarness::handle_event, and snapshot and describe onto
AgentHarness::frame_json and AgentHarness::describe_json, and it
implements the core’s Host trait by calling the world’s imports. A second
output, agent-component-lite, is the portable build without the NixOS-only
pieces.
An external flake gets it as cowboy.packages.<system>.agent-component. That
is how a host stays on one cowboy revision for both the component and the WIT
it was built from.
The starting point: templates/host
nix flake init -t github:dmadisetti/cowboy#host
nix develop --command cargo run
The template is the smallest embedder that runs: a flake taking cowboy as an
input, and one Rust binary (templates/host/src/main.rs) that instantiates
the component with wasmtime, answers the six imports in-process, calls
initialize with a tiny config map, submits one prompt, prints every
snapshot as a JSON line and exits. Its HTTP handler returns an error, timers
are delivered without waiting, and peer spawning is ignored. Replace those
stubs before using it for real agent work. Its devShell links wit/ to the cowboy
input’s specs/wit so the bindgen! bindings track the same revision as the
component. templates/host/README.md explains each export, each import, and
what to replace to put a real UI, HTTP client and clock around the loop.
The headless reference: montana
crates/montana/ is the host cowboy itself ships: the daemon under systemd on
NixOS, the process inside the container image, and the library the OCI
runtime (crates/runtime/) calls in-process. crates/montana/src/agent.rs implements the imports with threads —
exec and http run off the guest thread and post their results back into
one channel, timers are additive sleepers — and crates/montana/src/lib.rs
assembles the config map, resolves and preopens the agent’s directories at
their own paths, and starts the control socket. Frames stream out as ndjson
and intents come in the same way (specs/CONTROL.md); the optional A2A front
in crates/a2a/ is a protocol adapter over describe and handle-command
(specs/A2A.md). Read it when the template’s answer to a question is “a real
host does more here”.
Inside the component: the Host seam
Hosts that can link Rust do not need the component at all. The Zellij plugin
(crates/harness/) links cowboy-core directly and supplies a Host
implementation of its own. That trait, in crates/core/src/host.rs, is the
seam the WIT world mirrors. Commands and HTTP requests return results later
as events correlated by an opaque Ctx map. Timers, pane operations and peer
lifecycle requests also go through the host. AgentHarness::with_host takes the implementation; the entry point
after that is AgentHarness::on_load, in crates/core/src/harness.rs.
For a new asynchronous effect, update the Rust host trait and its adapters. If component hosts need it, update WIT and the component and Montana bindings as well. Terminal-only view hooks can remain Rust-side; headless adapters have no terminal view to update.
What is not an extension point
- Configuration keys are parsed in one place,
crates/core/src/config/; a new knob is a new key there and a new line in the launcher that emits it, not a host-specific side channel. - Tools and skills have configuration and module interfaces; see Plugin Architecture.
- A host’s boundary includes effect policy, WASI preopens and the enclosing process permissions, mounts and network controls. The Security Model describes the managed deployment; a new embedder must supply its own controls.
Concurrency Locking
The agent runs one session as an event loop. SessionLock
(crates/core/src/lock.rs) keeps LLM requests, tool execution and tool-output
summarization from interleaving, and queues user input that arrives while any
of them is in flight. Every hold on the lock is bounded: one that goes too long
without evidence of progress is reclaimed by a sweep the heartbeat runs.
What holds the lock
SessionLock tracks three kinds of hold:
llm: Option<Lease>— an LLM request is in flight.summarizing: Option<Lease>— a tool output is at the summary model.active_tools: HashMap<String, ActiveToolState>— tool executions keyed by their uniquecall_id, each with its ownstarted_atandtimeout.
is_locked() is true while any of them is set. at_tool_boundary() is the
inverse — no request, no summarization, no registered tool — and is the only
point at which queued input is replayed.
Leases and epochs
A Lease is a hold with a deadline: started_at, a timeout, and an epoch
naming this particular hold. The budgets live in lock::lease: LLM (300s)
covers one round trip with no intermediate signal, SUMMARIZING (180s) one
cheap-model summarization of a single tool output.
The deadline means “time since the work last moved”, not “time since
dispatch”. touch_llm() resets started_at on evidence the turn advanced — a
stream poll whose transcript grew, a compaction reply that starts its second
round trip — so a long but healthy streamed turn is never reaped, while a relay
that answers every poll with an unchanged buffer does not push the deadline
out.
try_acquire_llm(timeout) takes the LLM lease and returns its epoch, or None
when a request is already out. acquire_or_current_llm(timeout) returns the
live epoch when the lease is already held and takes it otherwise; compaction
uses it because it is reached both from inside a turn and from a free lock.
set_summarizing(timeout) takes the summarization lease and returns its epoch.
force_release_llm() and clear_summarizing() release them.
Epochs come from one monotonic counter shared by both leases, so an epoch
names exactly one hold for the life of the session. A dispatch stamps its
epoch into the request’s Ctx (dispatch::stamp_epoch, key EPOCH_KEY) and
the host echoes it back with the reply. Before any reply reaches a handler,
AgentHarness reads it (dispatch::ctx_epoch) and asks
epoch_is_current(epoch); a reply whose lease has been reaped or re-taken
belongs to a turn that has moved on and is dropped. A reply carrying no epoch
passes through — failing closed there would turn a host that drops the key
into a starved agent.
Tool tracking
register_tool(call_id, tool_name, timeout) records an ActiveToolState;
complete_tool(call_id) removes and returns it. Results are matched by
call_id, never by tool name, so concurrent calls to the same tool are tracked
independently. set_exec_kind(call_id, kind) attaches the execution-policy
classification made at dispatch, so a continuation of the same call is held to
the same policy. release_all_tools() drops every registration at once and
hands the states back: an interrupt settles the whole set, and the caller still
owes each one an abandoned-process record.
The sweep
sweep_expired() reclaims every hold whose deadline has passed — the LLM
lease, the summarization lease and timed-out tools alike — and returns them as
ExpiredLock values (Llm { waited, epoch }, Summarizing { waited, epoch },
Tool(ActiveToolState)). The leases come first so their handlers settle the
turn before a tool result handled afterwards can start the next request.
AgentHarness::sweep_session_lock (crates/core/src/heartbeat.rs) runs it at
the head of every timer fire, before any branch can return, and settles each
entry:
- an expired LLM lease drops any stranded relay stream, resets an active
status to
WaitingForInputand drains queued input; a late reply is then dropped by the epoch guard; - an expired summarization routes a synthetic 504 through
handle_summarization_response, which truncates the output it still holds, appends the tool result the model is waiting on and continues the turn; - a timed-out tool gets a synthetic tool result saying the harness stopped
waiting, naming the PID that was not cancelled —
Host::execis fire-and-forget, so a timeout can only stop waiting — and the process is tracked as abandoned.
hold_summary() renders what holds the lock and how long each hold has left;
submit logs it whenever input is queued, so a quiet agent says what it is
waiting on.
Input queueing and coalescing
submit (crates/core/src/intent.rs) calls queue_user_input(content) when
input_can_run_now() is false. boundary_reached() returns None unless
at_tool_boundary() holds, and otherwise drains the queue through
drain_ready_inputs():
- a single queued input is returned verbatim;
- several are coalesced into one message under a
[Multiple inputs received while processing:]header, joined by---.
execute_next_tool_call calls boundary_reached() after a tool result is
recorded, gated on the same input_can_run_now() predicate submit queues on,
so a rate-limit backoff — which holds no lease — cannot let queued input
through. The heartbeat’s tick also drains the queue (check_queued_inputs,
behind the same predicate), so input stranded by a lock-free window is picked
up without another turn event.
Tests
lock.rs covers LLM mutual exclusion, boundary detection, coalescing, tool
timeout removal, the combined locked-state predicate, the sweep reclaiming all
three kinds of hold in lease-first order, touch_llm keeping a live stream
from being reaped, and an epoch ceasing to be current once its lease is gone.
Result Summarization
Large tool outputs are stored on disk as recallable artifacts and replaced in the context window with a short stub. When summarization is enabled, the stub’s body is produced by a cheap “summary” model instead of blunt truncation.
Implementation: crates/core/src/tools/summarize.rs (classification,
prompts, request formatting, stubs, fallback truncation) integrated through
harness.rs (handle_command_result) and handlers.rs
(handle_summarization_response).
Threshold
A tool output is summarized when its length exceeds
DEFAULT_SUMMARIZE_THRESHOLD = 12000 characters; no launch-configuration key
changes that. Input sent to the summary model is capped at
MAX_SUMMARIZABLE_LENGTH = 50000 characters.
Summary models
The summary model’s provider is taken from the summary_model config and
resolved through the shared key chain. Default model per provider:
| Provider | Default model |
|---|---|
| Anthropic | claude-sonnet-4-20250514 |
| OpenAI | gpt-4.1 |
| OpenRouter | openai/gpt-4.1 |
Those three are the only summarizers there are (ResolvedProvider); there is no
Ollama or Codex summarizer. ollama — like any unrecognized provider name —
resolves to the OpenAI key and the OpenAI request shape, so a summary_model of
ollama:<model> does not reach a local Ollama server, it sends <model> to
OpenAI. codex resolves to the first key available in the order OpenAI,
Anthropic, OpenRouter, and because that is not a direct hit the configured model
name is dropped in favour of that provider’s default above. With no key at all,
no summarizer is built and outputs are truncated instead.
Requests use temperature: 0.3 and max_tokens: 2048 (or the
provider-appropriate completion-tokens parameter for OpenAI-compatible APIs).
Output classification
The tool name determines an output type, which selects a type-specific prompt
and fallback head/tail ratio (ToolOutputType::from_tool_name):
| Type | Tool names | Head ratio |
|---|---|---|
| FileContent | read, cat, __HASHLINE_READ__ | 0.7 |
| SearchResults | grep, rg, search, ast-grep | 0.5 |
| DirectoryListing | ls, find, fd, glob | 0.3 |
| CommandOutput | bash, sh (and any unknown tool) | 0.6 |
| StructuredData | nix-search, gh | 0.5 |
| WebContent | web-search, web-fetch | 0.5 |
Each type has its own prompt: command output preserves errors and exit codes verbatim, search results are grouped by file with line numbers, directory listings are grouped by kind, structured data extracts names and versions, and web content extracts main facts and quotes.
Code-file handling
For FileContent, SummarizationRequest::code_file splits the file into three
parts: a before-section and after-section (summarized) around a relevant range
that is reproduced exactly with hashlines. When no range is supplied the
relevant range defaults to the middle third of the file. The prompt instructs
the model to keep the surrounding summaries to one or two sentences and preserve
the hashline section verbatim.
Pipeline
- A tool result arrives in
handle_command_result. The full output is logged to the session’smessages.jsonlexactly once, regardless of what enters context. - If the output exceeds the threshold and the session is ready, it is stored as a disk artifact and kept in memory for UI expansion.
- If summarization is enabled, a
PendingSummaryis queued, the request is dispatched to the summary model, andset_summarizing(true)gates user input until the response returns. - In
handle_summarization_response, the lock is cleared, the provider-specific response is parsed, and the summary is wrapped in an artifact stub. On a parse failure or non-200 status,fallback_truncateis used instead.
If summarization is disabled, step 3 is skipped and the artifact stub is built
directly from fallback_truncate.
Artifact stubs and recall
format_artifact_stub wraps the summary with a header and a footer noting that
the full output is stored:
[Artifact <call_id> | <tool> | <bytes> bytes, <lines> lines]
<summary>
[Full output stored — use recall_artifact("<call_id>") to search or read more]
The agent retrieves the full output on demand through the recall_artifact
builtin tool (wire name __RECALL_ARTIFACT__), which reads or searches the
stored artifact by call_id.
Fallback truncation
fallback_truncate splits at line boundaries using the type’s head/tail ratio,
operates on character counts to stay UTF-8 safe, and inserts a
[...N chars omitted...] marker between the kept head and tail. It is used
whenever the summary model is disabled, unreachable, or returns an unparseable
response.
Inbound Message Reliability
A sender can receive no reply even though the agent started the work. A restart can also cause work or an external action to happen again. Before resubmitting a silent request, the operator must inspect the session, outbox, bridge outcome and any action the request could have caused.
Redis recovery retries delivery to dispatch, not to successful completion. The agent queues acknowledgment when it hands a message to the agent loop, without waiting for an answer. Once that acknowledgment lands, Redis will not recover an interrupted turn. If it does not land, reclaim can dispatch the message again. There is no exactly-once guarantee for completion or effects.
Session recovery is separate. On restart, the agent resumes its saved session, and an idle reprompt can act on the unanswered request again. The active message and reply-source binding lived only in process memory. The work can recur while the reply loses its source binding: a resumed answer may reach the session but no source outbox, leaving the original sender with silence.
| Observation | What it establishes |
|---|---|
| Service is running | The process is up; useful progress still needs inspection |
| No pending Redis entries | No entries await acknowledgment in that group |
| Acknowledgment landed | Dispatch was acknowledged, not completion |
| Reply in the agent outbox | A reply was queued for the bridge |
| Session resumed | Conversation state returned; source routing may not have |
There is no durable deduplication across agent restarts and no durable binding that reconnects a resumed turn to its source. Within a process, a bounded seen window suppresses repeat dispatch. Failed acknowledgments retry the acknowledgment first; exhausting that budget allows reclaim to offer the request again. External effects are not transactional with acknowledgment.
The operator procedure
Use the host-side operator identity to inspect missing or duplicate replies.
Check the agent with cowboy status and cowboy doctor first. A running
service does not prove that a request completed or reached its sender.
The identity. Not the agent’s. Every agent credential is scoped to
~<agent>:* and holds only the verbs the poll issues, which does not include
XPENDING — the agent has no use for it, and an agent that could enumerate its
own pending list still could not see another’s. modules/pubsub.nix generates
one identity for a person instead: operator, readable across the whole
keyspace, with no @admin (it cannot rewrite the ACL) and no @dangerous (no
CONFIG, FLUSHALL, KEYS, SHUTDOWN). Its URL is root-only, and Redis
binds inside the namespace, so every command below is ip netns exec:
sudo -i
. /run/cowboy/operator-redis.env # sets REDIS_URL
r() { ip netns exec cowboy-ns redis-cli -u "$REDIS_URL" "$@"; }
Three names come from the generated configuration rather than from memory. For
agent <agent> and source <source>, /etc/cowboy/<agent>/sources.json names
the stream; the group is cowboy; the consumer is harness-<source> — the
SOURCE, not the agent (RedisStreamsSource::new names it after the source it
polls; two agents on the same source carry the same consumer name on their
own separate streams):
stream=$(jq -r '.sources["<source>"].stream' /etc/cowboy/<agent>/sources.json)
group=cowboy
consumer=harness-<source>
1. Is anything stuck? The pending list is the only record of an entry the agent read and did not acknowledge:
r XPENDING "$stream" "$group"
A summary of 0 means the group has no pending entries. It does not prove
that every stream entry was read or that dispatched requests were answered.
2. Which entry, and how long has it been stuck? The extended form lists one line per entry — id, consumer, idle milliseconds, delivery count:
r XPENDING "$stream" "$group" - + 10
An entry whose idle time is under ten minutes is not yet eligible for reclaim.
It may be queued, or its reader may have stopped; age alone does not prove
progress. Reclaim waits until RECLAIM_IDLE_MS has passed. An entry idle for longer
than that whose delivery count is not rising is one no poll is reaching — check
that the agent is running at all before doing anything else.
3. What was in it?
r XRANGE "$stream" <id> <id>
4. Did the turn already have an effect? Ask this before step 5, because step 5 is the one action in this procedure that can do harm. The agent’s reply goes to its own outbox, and that stream is the record of it:
r XRANGE "<agent>:<source>:outbox" - + COUNT 20
An outbox entry shows a reply was queued, not that the external platform delivered it. Inspect bridge delivery state and the destination before resubmitting. Also check files, rebuild results and other requested actions.
Do not assume reclaim will discard a request that already had an effect. Deduplication is bounded and lives in process memory; after restart the same entry can dispatch again. Neither reclaim nor manual resubmission is an exactly-once operation. If work already happened, reconcile that outcome rather than blindly replaying the original request.
5. Resubmit only after reconciliation. If retrying is appropriate, submit a new entry. Keep in mind that the old pending entry can still be reclaimed. This minimal command submits content only; it does not restore the original sender or channel metadata and therefore does not promise a routed reply:
r XADD "$stream" '*' content '<the content from step 3>'
Implementation
What one poll does
poll_command emits three redis-cli calls in one sh -c:
XGROUP CREATE … 0-0 MKSTREAM, idempotent, so a stream that does not exist yet is not an error.XAUTOCLAIM <stream> <group> <consumer> <idle> <cursor> COUNT 10— the recovery step.XREADGROUP >moves an entry into the group’s Pending Entries List under the consumer that read it, and never offers a pending entry again —>means “entries no consumer has read”. So an entry a crashed run had taken is invisible to every subsequent>, and the reclaim is the only thing that hands it back. The consumer name is deliberately stable across restarts (RedisStreamsSourcedefaults it toharness-<source>, and the comment there says why): a fresh name per process would leave the whole previous pending list keyed to a consumer that will never ask for anything again. Stability keeps the list reachable; the reclaim is what reaches it.XREADGROUP GROUP … COUNT 10 STREAMS <stream> '>'— new entries.
Reclaim before read, so a message stranded by a restart is not queued behind
messages that just arrived. A marker line separates the two replies; both are
parsed by the same hand-rolled reader over redis-cli --no-raw output.
The scan carries its cursor. XAUTOCLAIM scans the pending list from a
cursor and answers with the one to resume from. A scan that spends its COUNT
on entries too young to claim returns an empty batch and a non-zero cursor;
restarting from 0-0 every poll re-walks that same head, and an orphan behind
a full scan’s worth of young entries is only reached once the head drains on
its own. RedisStreamsSource::reclaim_cursor keeps the returned cursor, which
is why poll_command and parse_poll_result take &mut self. Redis answers
0-0 when the scan wraps, so the walk returns to the head on its own.
An entry with no usable content is still emitted, with an empty one. It
has to be: an entry nothing returns is an entry nothing acknowledges, and since
the poll reclaims the pending list it would come back on every poll forever,
masking the ten entries behind it. The manager acknowledges and drops it.
Acknowledgment semantics
XACK answers with the number of entries it removed from the pending list.
The reply is examined, not the exit code alone: redis-cli prints a refused
command’s error text on stdout and exits zero, so an ACL that does not grant
XACK used to read as an acknowledgment that landed. The generated command
requires the reply to be a run of digits — zero included, which is the
idempotent case — and exits ACK_REPLY_NOT_A_COUNT otherwise.
The acknowledgment is queued at dispatch, not at completion.
process_pubsub_messages calls SourceManager::mark_processed on each message
as soon as the prompt has been handed to process_user_input, and the same
poll tick then calls send_pubsub_acks. Nothing waits for the turn to produce
an answer. Redis’s redelivery therefore covers exactly one window — from the
XREADGROUP that read the entry to the XACK that follows its dispatch — and
nothing after it. A process that dies mid-turn has already acknowledged the
entry it was working on; the pending list holds no record of it, and no reclaim
will bring it back. That is the deliberate trade: the alternative, acking at
completion, makes every crashed turn a guaranteed re-answer, including the ones
that had already sent their reply.
A failed acknowledgment retries the acknowledgment, not the request.
SourceManager::retry_ack queues the same XACK again, up to
MAX_ACK_ATTEMPTS. Only once that budget is spent does the manager forget the
entry, which hands it back to the reclaim. Forgetting on the first failure —
before the retry budget is spent — gives the reclaim a message the agent
has already answered, and the next poll starts a second turn on it.