Keyboard shortcuts

Press ← or → to navigate between chapters

Press S or / to search in the book

Press ? to show this help

Press Esc to hide this help

Cowboy

Persistent agents on infrastructure you control.

Cowboy runs persistent AI agents on infrastructure you control, using Nix to configure their tools, credentials, message access, and approved actions.

The strongest documented deployment is managed NixOS on Linux/x86-64. Local and ordinary Docker paths have different boundaries and hold credentials locally. The OCI runtime is a specialized optional route for Cowboy’s own payload.

Why run an agent this way?

A maintenance task can produce a reviewed change without giving the agent the credential that publishes it. The repository’s effects runner check exercises adding a documentation post to a Git repository: it describes a commit bundle, applies it, and verifies the remote branch. This is a recorded test workflow, not a claim about production adoption.

In a deployment with the effects bridge and Discord approvals configured, the agent requests publication. The publishing runner checks the bundle and produces a card naming the target, execution identity, exact SHA, changed files, and diffstat. An authorized human reacts ✅ to push exactly that commit or ❌ to reject it. The runner executes under the configured publishing identity, checks policy again, and refuses a branch that moved after review. The result appears in the agent’s effects inbox and the configured operator status channel. Gated effects explains the configuration.

Choose a deployment

GoalRouteBoundary
Evaluate in a terminalLocal CLI and ZellijInvoking user’s permissions; local credentials
Run a persistent managed serviceNixOS systemd service running MontanaPer-agent component, shared service hardening envelope, network namespace, credential proxy, Sheepdog, and broker controls
Run on an ordinary Docker hostSelf-contained imageDocker isolation; credentials in the container’s state volume
Integrate Cowboy’s payload with an OCI engineOptional Cowboy OCI runtimeBundle policy and configured cage; specialized command surface

The full deployment matrix includes host requirements and state locations. “Production path” means the intended supported managed configuration, not measured adoption or general runtime certification. Its agent daemon runs Montana directly, without the OCI runtime.

Network restrictions, filesystem access, tool policy, and approval gates are separate controls. HTTP policy restricts methods and destinations; it does not promise that every permitted request leaves external state unchanged. Data an agent can read can leave through permitted requests or replies. Read the security model before granting sensitive access.

Start here

Continue to Installation, which maintains release availability and source-build instructions. Then use First request and daily operation and Troubleshooting.

Configuration explains how to change the agent. Architecture explains the implementation and the cowboy:agent@0.2.0 component contract.

Deployment paths

Choose the boundary before installing. The strongest documented deployment is managed NixOS on Linux/x86-64. Local and ordinary Docker paths deliberately hold credentials; the optional OCI route runs Cowboy’s own payload. These paths do not form a ladder, and tool behavior depends on each path’s policy.

PathHost requirementsProcess identity and controlsCredentialsMutable statePrimary limitation
Local evaluationNix source build; Zellij for the UI; Linux or macOS with the dependenciesInvoking user; Zellij harness or native MontanaEnvironment, user/system key files, or readable agenix filesLocal XDG data/config directories and workspaceFile and shell tools have the user’s permissions; no managed proxy or service envelope
Managed NixOSNixOS, Linux/x86-64; configured operator and provider secretConfigured agent user; Montana under systemd, shared hardening envelope, network namespace, credential proxy, Sheepdog, broker ACLsProxy-readable secret files outside the agent; agent receives placeholdersAgent home, configured writable mounts, broker/bridge stateControls depend on configuration; readable data is not confidential from permitted egress
Ordinary DockerSource-built image; Docker on a matching Linux image architectureContainer root; ordinary Docker isolation; wizard uses Zellij, settings-based service uses MontanaMode 0600 key file in the configuration volumeConfiguration and workspace volumes; explicitly persist the data directoryNo managed credential proxy or Cowboy cage; settings-based CLI launch does not itself mount the default data directory
Optional OCI cageLinux/x86-64, Nix-built seed/rootfs, registered Cowboy runtime and resolver, Docker, companion imagesBundle-selected uid/gid; Cowboy runtime runs linked Montana with bundle syscall/filesystem policyOpenRouter credential in companion; placeholder in agentBundle-selected writable mounts; companion scratch on hostCurrent CLI companion supports OpenRouter only; Redis companion has no durable volume; not a general container runtime

Managed NixOS

Nix builds a per-agent component and declares the daemon, tools, proxy, message access, and service controls. The daemon launches Montana directly under systemd. It joins the configured Cowboy network namespace; the shared service envelope limits its writable world to the agent home and declared mounts. Sheepdog mediates tool execution. The daemon requires the credential proxy.

The module can also build an OCI bundle. That bundle is a separate deployment artifact, not an intermediate step in the systemd daemon launch. “Production” here describes the intended supported configuration, not adoption or runtime certification.

Local and ordinary Docker

Local UI sessions use the Zellij harness. Local headless sessions use Montana. Changing hosts does not add the managed security controls.

The ordinary image has two startup paths. The interactive wizard launches Zellij. A settings-based first boot generates and builds a Nix configuration, then its supervisor launches Montana on a Unix socket. Both keep real keys inside the container. The Docker walkthrough sets an explicit persistent data directory for the interactive path.

The NixOS module’s legacy container supervisor is a separate development artifact. It must not be confused with the managed systemd service or with the ordinary image’s settings-based supervisor.

OCI and embedding

The OCI runtime accepts Cowboy payloads and a narrow lifecycle command surface. It does not run arbitrary image commands. See the OCI contract before integrating it with an engine.

A custom host can implement cowboy:agent@0.2.0, but must supply and authorize its own effects and filesystem access. Component portability alone supplies no host security policy. Bare, user, and container are design classifications, not three runtime selectors.

Continue to Installation. The security model describes the conditions behind the controls in this table.

Installation

Choose a route in the deployment matrix first. Local, managed NixOS, ordinary Docker, and the optional OCI cage have different credential and process boundaries. Policy can change which tools work.

Release availability

Nothing has been published to a package registry. There is no published Python package or prebuilt Cowboy image. This page is the maintained release availability statement; all installation routes below build from source.

Start with a checkout and Nix with flakes enabled:

git clone https://github.com/dmadisetti/cowboy.git
cd cowboy

Keep the revision and lock file used for a deployment. The commands below run from this checkout unless they explicitly refer to your host configuration.

Local evaluation

Build the CLI and the repository’s Zellij variant:

nix build .#get-cowboy --out-link /tmp/cowboy-cli
nix build --impure --expr '
  let f = builtins.getFlake (toString ./.);
  in (import f.inputs.nixpkgs {
    system = builtins.currentSystem;
    overlays = [ f.overlays.zellij ];
  }).zellij' --out-link /tmp/cowboy-zellij
export PATH=/tmp/cowboy-cli/bin:/tmp/cowboy-zellij/bin:$PATH
read -rsp 'Anthropic API key: ' ANTHROPIC_API_KEY
export ANTHROPIC_API_KEY
cowboy --model anthropic:claude-opus-4-5-20251101

The Zellij overlay supplies the memory limit used by the harness. The key is in the launching environment. The CLI packages its harness WASM; a plain Python install from the checkout does not stage that artifact.

In Zellij, send:

Read README.md and summarize what this project does. Do not change files.

Expect tool activity and a model reply in the session, not a predetermined answer. File and shell tools run with your permissions. From another shell in the same environment, stop it with:

cowboy stop cowboy

For a local headless evaluation, build Montana and the component explicitly:

nix build .#montana --out-link /tmp/cowboy-montana
nix build .#agent-component-lite --out-link /tmp/cowboy-component
export MONTANA_BIN=/tmp/cowboy-montana/bin/montana
export MONTANA_WASM=/tmp/cowboy-component/share/cowboy/cowboy-agent.wasm
cowboy serve --model anthropic:claude-opus-4-5-20251101

This prints a Unix socket connection command. Connect to that socket and send one JSON line:

{"submit":"Say hello and describe your available tools."}

The socket emits display frames containing activity and the answer. Stop the unnamed headless instance with cowboy stop agent. This path still holds real keys and runs with your permissions; headless operation adds no confinement.

Managed NixOS

Use an existing Linux/x86-64 NixOS host configuration. Add Cowboy as a flake input, import the Cowboy and Home Manager modules into that host’s module list, and keep its lock file under version control. The fragment below assumes the host already imports agenix and has an operator account named alice; substitute your operator name. The encrypted key must be decryptable by this host.

# Add to the inputs of your host flake:
inputs.cowboy.url = "github:dmadisetti/cowboy";

# Include in the host's nixosSystem modules list:
modules = [
  ./configuration.nix
  cowboy.nixosModules.default
  cowboy.inputs.home-manager.nixosModules.home-manager
  ./cowboy-agent.nix
];

Save this module as cowboy-agent.nix beside the host flake:

{ lib, ... }:
{
  programs.fish.enable = true;
  services.cowboy.broker = "alice";
  services.cowboy.agents.dev = {
    enable = true;
    user = "dev";
    uid = 1338; # Choose an unused UID.
    model = "anthropic:claude-opus-4-5-20251101";
    daemon.enable = true;
    daemon.expose = 4312;
  };

  age.secrets.anthropic-key = {
    file = ./secrets/anthropic-key.age;
    owner = "cowboy-proxy";
  };
  services.cowboy.secretsProxy.domainMappings = lib.mkForce {
    "api.anthropic.com" = {
      secretPath = "/run/agenix/anthropic-key";
      headerName = "x-api-key";
    };
  };
}

The encrypted file is part of your host configuration; the plaintext key never belongs in Nix text or the store. Its runtime reference is the proxy mapping’s secret path. Restricting the mapping to this provider avoids provisioning unneeded provider and search credentials.

Build and activate from your host configuration directory, replacing HOST with its NixOS configuration name:

sudo nixos-rebuild switch --flake .#HOST
systemctl status cowboy-serve-dev --no-pager
cowboy connect 127.0.0.1:4312#/var/lib/cowboy/expose/dev.token

Run the connection command as the configured operator. It opens the Zellij client to the managed agent. Send “Say hello and describe your available tools.” Expect activity and a reply there; inspect service diagnostics with journalctl -u cowboy-serve-dev. The localhost control token grants control, including the ability to answer approval prompts. A read-only observer needs the separate observer endpoint and token.

The daemon runs Montana directly under systemd as dev, with its per-agent component, namespace, proxy, Sheepdog, and shared service envelope. It does not invoke the OCI runtime. To leave the daemon stopped:

sudo systemctl stop cowboy-serve-dev

A socket quit is insufficient: the unit restarts automatically. Set daemon.autostart = false and rebuild if it should remain off at boot.

Docker

Build on a Nix machine matching the target image architecture, then load the archive on the Docker host. The target host needs Docker, not Nix:

nix build .#docker-image --out-link /tmp/cowboy-image
docker load < /tmp/cowboy-image
docker run -it --name cowboy-eval \
  -e XDG_DATA_HOME=/etc/cowboy/data \
  -v cowboy-etc:/etc/cowboy \
  -v cowboy-workspace:/opt/workspace \
  cowboy:latest

For separate machines, transfer the image archive before loading it. On first boot the wizard asks for provider, model, and key. It writes the key with mode 0600 into the configuration volume, then opens the setup agent in Zellij. Tell it what work you want help with, your communication preferences, and when it should ask before acting. It saves the agreed profile to /opt/workspace/.cowboy/prompt.md; an otherwise empty workspace is normal. Choosing a role such as calendar assistant does not authorize or configure a calendar integration. Setup should state which capabilities work and which still need access, while preserving existing workspace files. The explicit data directory keeps session state in the named volume, including on later starts. Tools run as container root with ordinary Docker isolation; there is no managed credential proxy or Cowboy cage.

When the profile is saved, stop and resume from the Docker host. The normal agent loads that profile and resumes the conversation from persistent state:

docker stop cowboy-eval
docker start -ai cowboy-eval

Without a TTY or staged settings, a fresh image exits with “no configuration found.” For a settings-based headless service, the CLI can stage a bootstrap settings file and build a per-agent generation. That route is distinct from the interactive wizard. Its current launcher mounts config and workspace but does not set a persistent Montana data directory; account for that before relying on it for durable sessions. See daily operation.

Optional OCI cage

This specialized route requires Linux/x86-64, Docker, a source checkout, and a registered Cowboy runtime with its imageless resolver. Follow the registration instructions in crates/runtime/smoke.sh and the OCI contract. Registration changes Docker’s daemon configuration; it is not supplied by installing the CLI.

Once the runtime is registered, prepare the Cowboy payload and the placeholder image, then launch from the checkout with the CLI available:

nix build .#cowboy-agent-seed --out-link /tmp/cowboy-seed
nix build .#cowboy-agent-rootfs --no-link
tar -C /tmp/cowboy-seed/rootfs -c . | docker import - cowboy-imageless:latest
read -rsp 'OpenRouter API key: ' OPENROUTER_API_KEY
export OPENROUTER_API_KEY
cowboy serve demo --sandbox cage --expose 4210

The companion currently injects only OpenRouter credentials; another provider’s key will not work here. It holds the real key and exposes egress through Unix sockets. Docker runs the payload with the Cowboy runtime and networking disabled; the runtime launches linked Montana under the bundle’s policy.

Use the printed cowboy connect command and send the same hello request. Expect display frames in the client. The component and model come from the seed configuration. Stop the cage and its companions with cowboy stop demo. The companion Redis has no persistent volume; this is not the managed NixOS service’s persistence contract.

Continue with First request and daily operation.

First request and daily operation

Use one of the complete routes in Installation. Start with a request whose result you can check, such as “Read README.md and summarize what this project does. Do not change files.” Check the tool results and answer in the client. A successful launch alone says nothing about model authentication or useful work.

Inventory, status, and health

cowboy list
cowboy status
cowboy doctor

These answer different questions. The list shows registered agents that can run. Status shows runtime instances and their last reported activity. Doctor checks whether services are up, the event loop responds, and inbox consumers are still reading. A local session need not have a registry entry, and doctor cannot verify every deployment from every account. An unknown result is not a pass. Use Troubleshooting to interpret it.

For a named managed agent:

cowboy status dev
cowboy doctor dev --json
journalctl -u cowboy-serve-dev --since today

cowboy logs reads the launch log of a backgrounded local serve process; it is not a general transcript or journal reader. For example, cowboy logs agent reads the unnamed headless instance’s log when it was started with --daemon. Use the system journal for managed units, Docker logs for containers, and the client/session history for conversation content.

Review external actions

Approvals are configured per bridge or effect. For a gated Git push, review the runner’s target, execution identity, commit SHA, changed files, and diffstat. The agent’s prose is not the runner’s description. An authorized approver chooses ✅ to push exactly that SHA or ❌ to reject it. An expired request needs a new review; a branch that moved requires a newly prepared request.

The operation result returns to the agent’s inbox and the configured status channel. Keep the request ID and SHA if delivery or execution is uncertain. An accepted message or approval is not proof that the push or rebuild finished. See approvals and gated effects.

Restart and recover

Managed daemons restart after exit, including clean socket quits, with a five-second delay and a start-rate limit. A changed per-agent component also triggers a restart on NixOS activation. Use systemd for an explicit restart:

sudo systemctl restart cowboy-serve-dev
cowboy doctor dev

A deliberate systemd stop leaves the daemon down. Ordinary containers in the installation example have no Docker restart policy; start them explicitly. For local headless and cage launches, retain the original launch command: cowboy restart cannot reconstruct their transient socket, sandbox, and exposure flags. It stops those instances and asks you to relaunch. For a settings-based Docker agent it can reuse the declared volume.

After a crash, a sender may receive no answer even though a request was accepted, or see a duplicate after redelivery. Resumed work can also lose the original reply destination; inspect state and external outcomes before resubmitting. Message reliability gives the exact recovery contract.

Back up mutable state

Stop the agent and its writers before copying state. Keep ownership and modes, especially for keys. A Nix closure rebuilds software; it does not restore conversations, working files, credentials, or pending approvals.

  • Local: preserve the workspace and Cowboy’s XDG data/config directories (normally under ~/.local/share and ~/.config).
  • Managed: preserve the configured agent home and writable mounts, plus the enabled broker/bridge services’ persistent state and secret sources. Include Redis persistence when recovering queued messages and approvals matters.
  • Docker walkthrough: preserve both named volumes, including the explicit data directory under the configuration volume. Container removal does not remove these named volumes.
  • Settings-based Docker CLI: config and workspace are under ~/.local/share/cowboy/agents/<name>/ by default. The current supervisor leaves Montana’s default data directory in the container layer. Preserve it separately before replacing the container, or configure a data directory in a mounted volume in the generated supervisor. The config/workspace mounts alone are not a complete session backup.
  • OCI: preserve the writable mounts you declared in the bundle; do not treat runtime scratch or the CLI companion’s Redis as a durable backup.

Upgrade and rebuild

Keep the old revision, lock file, and a state backup until the new deployment has answered a test request. Build the new CLI or image from the chosen source revision. On NixOS, update the host’s Cowboy input and rebuild the host, then check doctor and the first real request.

For the Docker walkthrough, stop and remove the old container after backing up, load the new image, and repeat the run command with the same named volumes. A settings-based container also has a generated flake and result link under its configuration volume. Edit that persistent configuration and run cowboy bake inside the container to build a generation; restarting with an existing result link does not itself rebuild it. A new base image alone does not replace that existing generation.

An approval-gated rebuild pins the main commit and resolved input overrides before asking for approval. The result reports when a successful build produced the same system or home generation. Treat that as no deployment change, not evidence that the requested behavior reached the running agent. See rebuild diagnostics.

Stop and remove

Stop local or CLI-managed agents with cowboy stop <name>. For a managed daemon, stop its systemd unit, disable the agent in the NixOS configuration, and rebuild. For the Docker walkthrough, stop and remove the named container. Keep backups, then deliberately remove only the state and volumes you no longer need. Removing software or disabling a service is separate from erasing its history and keys. Revoke provider or publishing credentials when retiring the deployment.

Continue with Configuration to change models, tools, message access, and service settings.

Troubleshooting

Start with cowboy doctor. A service can be up while the agent has stopped reading messages.

cowboy list
cowboy status
cowboy doctor
cowboy doctor dev --json

The list is the registered inventory. Status observes running instances and last reported activity. Doctor separates three observations:

ObservationWhat it establishesWhat it does not establish
Process/service existsThe host still has a running processThe event loop is making progress
Event loop respondsA diagnostic request produced a new frameInbox polling still runs
Inbox consumer is activeRedis records recent readsThe model finished a turn or a reply reached its destination

A replayed status frame can outlive a stalled loop. Doctor asks for a new one. Its attention check examines inbox consumer activity independently.

Checks report ok, fail, or unknown. Unknown means the probe could not establish an answer, for example because the caller cannot read an agent’s Redis socket or the deployment has no supported probe. Read the reason; do not translate it into healthy. On a managed host, rerun as an authorized operator with the required privilege if the reason is access. Do not broaden socket permissions to make a diagnostic green.

Find the right log

# Managed service:
systemctl status cowboy-serve-dev --no-pager
journalctl -u cowboy-serve-dev --since today
systemctl list-units 'cowboy-*'

# Backgrounded local serve (default headless name):
cowboy logs agent

# Docker walkthrough:
docker logs cowboy-eval

Inspect the first failed service and its dependencies. The managed daemon requires the namespace and credential proxy; it does not launch an OCI bundle. A cage startup failure instead belongs to Docker, its runtime registration, the seed/rootfs, bundle policy, or companion services.

The CLI or Zellij is missing

Use the source-build installation. Select the CLI package explicitly; the default flake package is not the Python launcher. A Python source install alone does not bundle the harness WASM. Zellij is required for interactive sessions and the connect client, not for Montana’s headless service.

The client opens but the model fails

Check the provider/model specification, the credential source for this route, and the upstream response. Local and ordinary Docker paths hold real keys; managed agents use placeholders and need the proxy’s secret path to be readable by its service user. User and system keys files accessible to group or others are ignored. See secret discovery.

On a managed daemon, a runtime provider switch also needs provider access in the Nix configuration. A proxy credential alone does not supply the component’s provider key placeholder. Declare the provider through a model option or extraProviders and rebuild.

Record the provider, model, HTTP status, and request ID without copying keys. A method or destination denial calls for reviewing the configured policy, not bypassing the proxy.

A tool is denied

Identify the layer from its result: tool approval, filesystem permissions, Sheepdog syscall policy, service hardening, proxy methods/destinations, or OCI policy on the cage route. Compare the operation to that deployment’s declared access. A successful local run is not evidence that the managed policy should permit it.

An approval or reply seems stuck

Follow the request through ingest, the agent inbox, the agent turn, and the outbox. For Discord, the receipt reaction establishes that the bridge queued the message, not that the agent answered. Check doctor for continued inbox consumption and inspect whether an approval is pending, rejected, expired, or superseded. Only configured approvers can resolve a gated Discord request.

For Git effects, preserve the request ID and SHA. The approval must match the runner’s description and the applied SHA. A branch moving after review causes a refusal; fetch and prepare a fresh request instead of treating the old approval as permission for another commit. A push can succeed while its verification fetch fails: inspect the remote before retrying.

After interrupted work, a request may have no reply, be redelivered, or resume without its original reply destination. Inspect external results before asking for the same side effect again. See message reliability.

A rebuild succeeds but nothing changes

The rebuild service compares the built system closure with the running system, or the new home activation path with the prior one when it can read that state. An identical result gets an explicit no-change note. A successful command can therefore mean no new deployment.

First check the rebuild type: a home rebuild cannot deploy system services, networking, or system packages. Check that the approved commit contains the change, that it reached the tracked branch, and that the flake imports the edited file. Check resolved input overrides in the approval and result. Then check the deployed generation and whether the daemon restarted onto the new component. Doctor compares the generated agent configuration timestamp with the daemon start time to flag possible configuration drift. If the rebuild could not establish a prior home generation, absence of a no-change note is not proof of change.

For settings-based containers, an existing result link is reused at startup. Build the edited generated configuration with cowboy bake inside the container before restarting; merely editing JSON or replacing the base image does not necessarily replace that generation.

Restart without losing the evidence

Copy mutable state before repair. Managed daemons restart after exit, but a live stalled process need not exit; record doctor and journal output before an explicit systemd restart. Recheck health and send a small request afterward. See daily operation for restart limits, backup, and removal.

Report a useful bug

Include the source revision, OS/architecture, deployment route, redacted command or Nix configuration, expected and observed behavior, doctor output, and a bounded log excerpt. For bridge/effect problems include the request ID and commit SHA. Report at GitHub issues.

Configuration

CLI launch settings, persistent JSON, generated NixOS configuration, and secret discovery are different inputs. A provider key is not another level in ordinary settings precedence.

Local settings

The launcher merges JSON objects in this order, with later files winning:

  1. /etc/cowboy/config.json
  2. ~/.config/cowboy/config.json (under the configured XDG config directory)
  3. .cowboy/config.json in the current directory

Missing, unreadable, malformed, and non-object files contribute no settings. Explicit CLI options override the corresponding ordinary settings.

For provider, main model, and heartbeat, merged files override generated agent JSON, which overrides built-in defaults. For secondary models and memory/rerank settings, generated agent JSON overrides merged files; supported CLI options still win. Do not assume one precedence order covers every field. The declared systemd daemon takes its arguments from NixOS configuration directly, rather than reading the operator’s local launch defaults.

A local config can contain:

{
  "provider": "anthropic",
  "model": "claude-opus-4-5-20251101",
  "log_level": "info"
}

Common settings subset

This is a common subset, not the entire recognized surface. The implementation inventory is crates/contracts/src/keys.rs; a host or launcher need not expose every component key as a CLI option.

KeyMeaning
providerProvider for the main model
modelBare model ID in local JSON; provider comes from the separate key
summary_modelProvider:model specification for summarization
compact_modelProvider:model specification for compaction
subagent_modelProvider:model specification for sub-agents
vision_modelNative vision or a separate provider:model specification
memory_backendMemory search backend
heartbeatLauncher heartbeat interval, in seconds
log_levelLogging verbosity

The component key for heartbeat is heartbeat_interval; the launcher’s JSON setting and flag use heartbeat and --heartbeat. Model availability and provider access are separate: selecting a model does not provision credentials. Use cowboy models to inspect the bundled catalog and each command’s --help for its options.

Secrets

Outside proxy deployments, discovery checks each key in this order and uses the first nonempty value:

  1. Provider environment variable, such as ANTHROPIC_API_KEY.
  2. User keys file, ~/.config/cowboy/keys.json.
  3. System keys file, /etc/cowboy/keys.json.
  4. The provider’s agenix file, if readable.

Keys files map component key names, such as anthropic_api_key, to secret values. They must have mode 0600: files accessible to group or other users are skipped with a warning. Secret files and environment variables remain real credentials in local and ordinary Docker deployments.

Managed daemons receive proxy placeholders. The actual key belongs at the secret path named by the proxy domain mapping, readable by the proxy service user. The managed installation shows an agenix declaration with that ownership.

First-boot settings

A --settings file for cowboy init or cowboy serve is a validated bootstrap input, not the general config file. It accepts a key reference through api_key_env or api_key_file; a literal api_key is rejected. For example:

{
  "agent_name": "myagent",
  "posture": "container",
  "provider": "anthropic",
  "model": "claude-opus-4-5-20251101",
  "api_key_env": "ANTHROPIC_API_KEY"
}

The host resolves the reference and writes the key into the agent’s configuration volume with mode 0600. A bootstrapped volume rejects restaging; edit its persistent configuration instead. See daily operation for generation and backup considerations.

NixOS module

Import the default NixOS module and enable individual agents under services.cowboy.agents. There is no top-level service enable switch. Keep agent usernames equal to their agent names when using Redis ACLs, and give each enabled agent a unique UID.

The main model option takes a provider:model specification. Common per-agent options include homeDirectory, prompts.system, mounts, daemon.enable, and daemon.expose. The latter provides an authenticated control endpoint; daemon.observe provides a separate read-only endpoint and credential.

Nix generates per-agent launch configuration and components, and message-source configuration for agents and bridges. Edit the Nix declaration, not the generated files under /etc/cowboy. The full per-agent schema is in modules/options/user.nix.

Managed provider access is provisioned from the model options. To keep another keyed provider available for a runtime model switch, declare it in extraProviders and provision its proxy secret. This grants provider access; it does not select the main model.

Continue with tools and skills and bridge integration.

Approvals & Outbox


title: Outbox Approval Protocol tags: [pubsub, approval, discord, email]

Outbox Approval Protocol

Outbox services can hold an outbound message for human approval before sending it. The agent never sends directly; it writes to an outbox stream, and a broker service decides whether to forward the message, optionally gating it behind a human reaction.

Source: pkgs/bridge/base.py (OutboxService), pkgs/bridge/ingest.py (IngestService), with per-platform implementations in pkgs/discord/ and pkgs/email/.

See also Security Model.

Transport

Messages move over Redis Streams. Each bridge {name} has an inbox stream ({name}:inbox) and an outbox stream ({name}:outbox). OutboxService reads its outbox via a consumer group ({name}-outbox / worker {name}-worker) and acknowledges each entry after handling it. Approval state lives in plain Redis hashes so any service can poll it.

Flow

When approval_required is set, an outbound message is held and routed through a notification channel:

  1. The agent writes a message to its outbox — {agent}:{source}:outbox under the per-agent ACL (the flat {source}:outbox on a shared Redis).
  2. OutboxService.process_one() sees approval_required and the message has no approval_id, so it parks the entry (_resume_or_start_approval):
    • creates approval:{uuid} with status=pending, source, channel_id, a content_preview (first 200 chars), created_at and expires_at
    • writes a notification to the configured notify outbox ({approval_notify}:outbox) carrying the approval_id
    • returns without acking, so the rest of the stream keeps moving
  3. The notify service (e.g. Discord) sends the notification. Because the message carries an approval_id and the send returns an external_id, the base loop records the mapping in approval:{source}_map (external id → approval id).
  4. A human reacts to the notification.
  5. The notify service’s ingest handler calls OutboxService.resolve_approval(), which looks up the approval id from the map and, if the record is still pending and inside its window, sets status to approved or rejected plus the approver. A late answer — a record already decided, past expires_at, or already settled and deleted — is ignored; the map entry is dropped either way.
  6. The original outbox service checks its parked entries once per loop (poll_awaiting, via check_approval, which reads the deadline before the decision); it sends the message on approved, or acks it and tells the agent on its inbox that it was rejected or expired.

State machine

approval:{uuid} = {
  status:          pending | approved | rejected | expired
  source:          email | discord | ...
  channel_id:      destination (email address, channel id, ...)
  content_preview: first 200 chars of the message body
  created_at:      unix timestamp
  expires_at:      unix timestamp; past it the record reads as expired
  approver:        who resolved it (set on resolution)
}
pending --+-- approve --> approved --> send
          +-- reject  --> rejected --> drop, notify the agent
          +-- timeout --> expired  --> drop, notify the agent

expired is never written: it is what check_approval answers for a record past expires_at, whatever its status field says. poll_awaiting deletes the approval hash when it settles the entry.

Redis keys

approval:{uuid}         # HASH — approval state
approval:{source}_map   # HASH — external message id -> approval id

The requesting service writes and polls approval:{uuid}. The notify service owns approval:{source}_map. Bridges share only the key format — Discord and email do not know about each other.

Implementation

The park-and-poll flow (process_one(), poll_awaiting()), the static resolve_approval(), and the auto-tracking of the map all live in the shared OutboxService. A per-platform bridge only needs to:

  1. Return SendResult(ok=True, external_id=...) from its send().
  2. Call OutboxService.resolve_approval(pubsub, source, external_id, status, approver) from its ingest handler when a reaction or reply arrives (IngestService.try_resolve_approval() wraps this).

systemd services

Each bridge {name} is run as three systemd services: cowboy-{name}-ingest, cowboy-{name}-outbox, and cowboy-{name}-ping.

Configuration

Approval is configured per bridge under services.cowboy.bridges.<name>.approval (and equivalently on services.cowboy.pubsub.sources.<name>.approval):

services.cowboy.bridges.discord.approval = {
  required = false;       # hold outbound messages for manual approval
  notify = "discord";     # which outbox to send approval notifications to
  notify_channel = "";    # channel id within that outbox
  timeout = 3600;         # auto-reject after N seconds (0 = no timeout)
};

If required is true, notify_channel must be set. The Redis ACL enforcement option keeps the agent restricted to stream commands so it cannot touch the approval hashes directly.

Two bridges use this protocol for something other than a message. rebuild gates a system rebuild, and effects gates a declared action on the world (a git push, for instance) — see Gated Effects. For effects, approval.required is read-only true: there is no unapproved mode, and the approval is bound to a specific commit sha rather than to the request text, so what the operator approved is what runs, byte for byte.

Adding a new approval channel

To approve via something other than Discord:

  1. Have the new service’s outbox return external_id from send().
  2. Have its ingest call OutboxService.resolve_approval(...).
  3. Set approval.notify = "<service>" on the bridge that needs approval.

No changes to OutboxService or the requesting service are required.

Gated Effects

An effect is an action on the world that an agent may propose but must not perform: pushing to a branch someone else consumes, publishing, deploying. services.cowboy.pubsub.effects gives each declared effect a fixed shape — the agent names a commit, a human sees exactly that commit on an approval card, and after ✅ the host applies exactly that commit as a user of the operator’s choosing. Nothing the agent says on the card is trusted; nothing the agent can do between the card and the push changes what gets pushed.

It reuses the Outbox Approval Protocol unchanged. What is new is the seam on the far side of the gate: a bridge with no privilege of its own starts a static systemd template unit through polkit, and that unit — running as the declared runAs user inside the standard sandbox — does the work.

The flow

agent                      effects bridge                     host
-----                      --------------                     ----
commit in workspace
effects-request tool ───▶  effects:outbox entry
  (git bundle into        pre_send:
   the handoff dir)         systemctl start --wait
                              cowboy-effect-<name>@describe-<sha>  ──▶ unit (User=runAs):
                                                                      unbundle, fetch, policy,
                                                                      write results/describe-<sha>.json
                            read the result; refuse ⇒ ack + message
                            ok ⇒ pin sha/base/description
                          approval card → notify channel
                                       ⏳ … human reacts ✅
                          send():
                            guard pinned sha == requested sha
                            mark_effect
                            systemctl start --wait
                              cowboy-effect-<name>@apply-<sha>  ─────▶ unit: re-fetch, re-run policy,
                                                                      push --force-with-lease,
                                                                      verify, write apply-<sha>.json
                            result → agent inbox [effects <id>] …
                                   → notify channel (untagged copy)

Two runs of the same program, addressed by the same 40-hex sha: describe before the card, apply after it. The sha in the unit instance name is the whole of what the bridge can choose; polkit permits only that one action, on only those unit names, for only the bridge user.

Declaring an effect

services.cowboy.pubsub.effects = {
  enable = true;
  approval.notify_channel = "<discord channel id>";   # required; there is no unapproved mode
  declared.blog = {
    runAs = "<user>";                                  # the unit's User=; never root, never the bridge
    gitPush = {
      remoteUrl = "git@github.com:me/blog.git";        # ssh://, git@, or file:/// — never https
      branch = "main";
      webUrl = "https://github.com/me/blog";           # optional; result messages link the commit
      allowedPaths = [ "docs/posts/" ];                # directories (or exact files) every commit must stay inside
      maxCommits = 3;
      knownHosts = "github.com ssh-ed25519 AAAA…";     # required for ssh remotes
    };
    credentials.sshkey = "/home/<user>/.ssh/keys/github"; # loaded with LoadCredential=, readable by root only
  };
};

approval.required is read-only true. An effects bridge that starts without approval configured exits at authenticate(), and the module refuses to evaluate without a notify channel or with the notify bridge disabled — the failure mode where a “gated” effect quietly evaluates to an ungated one is the one this module exists to close.

Each declared effect produces:

  • a group cowboy-effect-<name> whose members are the agents plus runAs (never the bridge), and a setgid handoff directory /var/lib/cowboy-handoff/<name> (2770) where the agent’s bundles land;
  • a state directory /var/lib/cowboy-effect-<name> owned by runAs holding the mirror repo.git (0700), results/, and the lock;
  • a template unit cowboy-effect-<name>@.service with User=runAs, NoNewPrivileges, ProtectSystem=strict, ProtectHome, an empty capability set, IPAddressDeny outside the remote, and the docker socket inaccessible;
  • one polkit rule allowing cowboy-bridge-effects to start (and only start) cowboy-effect-<name>@(describe|apply)-<40 hex>.service;
  • an effects-request tool visible to the effect’s agents.

Set runner instead of gitPush to supply your own program. It receives the instance name (describe-<sha> or apply-<sha>) as $1, runs as runAs with STATE_DIRECTORY set, and must write $STATE_DIRECTORY/results/<verb>-<sha>.json with at least {"ok": bool, "message": str}; a describe result should also carry what the card shows (base, count, commits, files, diffstat, policy).

What the agent does

The effects-request tool takes effect, sha, repo, and content. It verifies the sha is a commit in repo and that refs/remotes/origin/<branch> exists there, writes a git bundle of sha ^origin/<branch> into the handoff directory (size-capped), and files one entry on effects:outbox. It returns {"status": "pending_approval", "request_id": …} and nothing has run.

The outcome arrives later in the agent’s inbox as a message starting [effects <request_id>]. A refusal — the branch moved, a file outside allowedPaths, a merge commit, too many commits, a symlink — is text, not a retry: the agent rebases and files a new sha, which is a new approval. The policy is checked on every commit in the range, not on the net diff, so an add-then-revert pair cannot smuggle a path into history that the card never showed. A newer request for the same effect from the same agent supersedes an older one still waiting; a refused request supersedes nothing.

What the human sees

Approval needed [effects/blog]: push to git@github.com:me/blog.git main (as <user>)
Requested by agent: `publish monix — reviewed, no blockers`
Commit 3f9c2a1e…  (fast-forward from 096ab9d, 1 commit, policy ok)
  3f9c2a1 blog: publish monix (agent-gated) (agent, 2026-08-25)
Files:
  A  docs/posts/2026-08-25-monix.md
 1 file changed, 118 insertions(+)
Inspect: git -C /var/lib/cowboy-effect-blog/repo.git show 3f9c2a1
React ✅ to push exactly this commit, ❌ to reject. Expires in 60 min; if the branch moves first the push refuses itself.

Every line but the quoted Requested by one comes from the describe run, which read the bundle as runAs, not from the agent. The agent’s own text is a single quoted line, capped, with backticks neutralised, so it cannot imitate the trusted lines. The Inspect: path is a real mirror on the host for a second look before reacting.

What can still go wrong

  • The branch moves. apply refetches and requires the branch to be at the base describe recorded; the push itself is --force-with-lease against that base. Either check failing is a refusal with the new head named.
  • The unit is interrupted. The bridge marks the entry phase=effect before starting apply, so a restart mid-push reports uncertainty instead of re-running. on_uncertain reads results/applied-<sha>.json and says whether the push in fact completed.
  • A stale result. The bridge treats a result file older than the unit it just started as absent — a describe that failed to run (polkit said no, say) cannot be answered by the previous run’s file.
  • runAs is a login user. The unit inherits that user’s DB groups (the module warns). The sandbox is what bounds them; a dedicated uid with a deploy key is the better end state, and is a one-line change to runAs and credentials.sshkey.

See Approvals & Outbox for the protocol underneath and specs/THREAT-MODEL.md §8 for the residual authorities this adds.

Rebuilds and generated environments

The rebuild bridge is separate from a declared Git effect. Enable it through services.cowboy.pubsub.rebuild and configure its repository and approval settings. The rebuild-request tool queues a system or home rebuild and returns a request ID. A queued or pending-approval result is not a completed rebuild. The later outcome arrives in the agent inbox tagged [rebuild <request_id>].

A system request builds the configured remote revision with nixos-rebuild boot, then starts activation in a separate systemd unit. A home request builds and activates the requesting agent’s home-manager profile. A home rebuild cannot apply system service or networking changes. Local unpushed edits are not the remote revision the bridge builds.

For system requests, a success message means the build finished and activation was launched; it does not prove the asynchronous switch completed. Inspect the activation unit and running service before calling the change live:

sudo systemctl status cowboy-activate.service
sudo journalctl -u cowboy-activate.service
cowboy doctor

The bridge compares the built system with the previously running generation, or the built home profile with the previous home generation. When it can prove they are equal, its result says the rebuild deployed no change. Check the request type, pushed revision and selected flake inputs before retrying. An unreadable generation is unknown, not evidence of a no-op. After an interrupted activation, inspect the running system and the bridge’s status record at /var/lib/cowboy-rebuild/last-rebuild.json before submitting again.

Ordinary Docker also has implemented generation machinery: settings can scaffold a flake in the persistent configuration volume, bake its container output and launch the resulting supervisor. The host records generation metadata for cowboy list; that record is not a liveness check. Configuration and workspace volumes are mutable state to preserve across replacement. See Deployment Paths for launch setup.

These mechanisms generate and apply environments. They do not implement an autonomous rollback or pull-request review workflow.

AI Governance

Cowboy constrains what an agent can do through several independent mechanisms. They overlap deliberately: a command rejected by one layer is not relied upon to be caught by another.

LayerWhereWhat it does
Approvalsbridge / Redishold outbound messages for a human reaction
Egress allowlistsecrets proxyrestrict HTTP methods and destinations, inject credentials
Sheepdogseccomp sandboxenforce file/network/exec rules at the syscall level

This page summarizes how they are configured. There is no runtime rule engine or FilterAction-style API in the harness — governance is the sum of the mechanisms below. Command, path, and syscall enforcement is done by sheepdog at the kernel boundary (see below), not by a separate string-matching filter layer.

Approvals (human in the loop)

For outbound messages that should not be sent autonomously, the bridge approval protocol holds a message until a human reacts to a notification. Configure it per bridge:

services.cowboy.bridges.discord.approval = {
  required = true;
  notify = "discord";
  notify_channel = "<channel-id>";
  timeout = 3600;
};

Approval state is tracked in Redis hashes and resolved from human reactions. See Approvals & Outbox for the full protocol.

Egress allowlist (secrets proxy)

In the managed proxy deployment, network rules route HTTP through the proxy and the agent uses placeholder provider credentials. Local and ordinary Docker launches hold credentials instead. With egress control enabled, GET, HEAD, OPTIONS and TRACE pass the method gate; every other method needs an allowed destination or receives HTTP 403. Allowed requests can still change external state or disclose readable data. See Security Model.

services.cowboy.secretsProxy = {
  enable = true;
  domainMappings = {
    "api.anthropic.com" = {
      secretPath = "/run/agenix/anthropic-key";
      headerName = "x-api-key";
    };
  };
  # Extra write-allowed domains. Domains in domainMappings are implicitly
  # write-allowed.
  allowedWriteDomains = [ "github.com" "api.github.com" "*.githubusercontent.com" ];
};

The allowlist enforcement lives in the mitmproxy addon (proxy/addon.py); write methods, the allowed-domain check, and wildcard matching are implemented there. See Security Model.

Sheepdog (seccomp sandbox)

Sheepdog enforces file, network, and exec rules at the syscall level rather than by string matching. Rules are verb-granular — Bash, Read, Edit, Create, Delete, and Connect — and resolve to allow or deny, with optional runtime-granted exceptions (lazy permissions) taking precedence over baked-in denies.

services.cowboy.sheepdog = {
  enable = true;
  lazyPerms = true;   # allow runtime permission grants
};

services.cowboy.agents.<name>.sheepdog = {
  deny  = [ "Connect(0.0.0.0/0)" ];
  allow = [ "Read(/home/*/workspace/**)" "Edit(/home/*/workspace/**)" ];
  blockedSyscalls = [ /* ... */ ];
  readonlyPaths = [ /* ... */ ];
  maskedPaths   = [ /* ... */ ];
};

Sheepdog is Linux-only. See crates/sheepdog/src/policy.rs and modules/options/sheepdog.nix.

See also

Plan state is advisory

The agent can call update_plan_state to record an objective, tasks, blockers and the next step. Each task is pending, in progress or completed. Updates replace the supplied fields, keep at most one task in progress and record an update timestamp.

The state is saved as plan_state.json in the session directory and loaded on resume. A non-empty plan appears in the model’s transient context packet, including up to eight tasks. It does not need to be repeated in the transcript.

This is a memory aid, not an approval gate. Plan state does not restrict the tool list, impose a read-only planning phase or enforce that actions follow the listed tasks. Use bridge approvals and host policy for enforced controls.

Memory System

Persistent, cross-session memory for the agent. Notes are stored as Markdown files and surfaced back into the model’s context via a search step before the agent acts.

Source: crates/core/src/memory.rs. Tools are registered in crates/core/src/tools/builtin.rs.

Backends

Memory is pluggable behind the MemoryBackend trait. Two backends exist:

  • ZkBackend — zettelkasten built on the zk CLI. Keyword-based full-text search. Always available, including lite builds. This is the default for lite.
  • QmdBackend — hybrid BM25 + vector search via the qmd CLI, using zk for note authoring. Available only in ranch (non-lite) builds, where it is the default (memory_backend = "qmd").

The backend is selected by the memory_backend config key ("qmd" or "zk"). qmd_min_score (default 0.4) sets the relevance cutoff for the qmd backend. If a lite build is asked for qmd, it falls back to zk.

Memory backends return shell command strings. Core asks its host to execute them and parses stdout into notes. This works through either Zellij or Montana; the required memory CLI must be available in the host’s execution environment.

Tools

The agent interacts with memory through two built-in tools:

  • __MEMORY_SAVE__ — write a note.
  • __MEMORY_SEARCH__ — search notes.

Notes

A parsed note (MemoryNote) has:

FieldMeaning
ididentifier, usually the filename without extension
titletitle from the note’s frontmatter
pathfull path to the note file
tagstags associated with the note
bodynote body (may be truncated)
scorerelevance score 0.0–1.0; only set by QmdBackend

Directory layout

Notes live under <cowboy_dir>/memory/, a zettelkasten managed by zk:

<cowboy_dir>/memory/
  .zk/        # zk configuration and index
  daily/      # YYYY-MM-DD.md journals
  facts/      # atomic knowledge notes
  decisions/  # decision records
  templates/  # zk note templates

<cowboy_dir> is $HOME on a ranch install and $XDG_DATA_HOME/cowboy (default ~/.local/share/cowboy) in lite mode.

Pre-tool retrieval

Memory search is wired in as a retrieval step before the agent runs tools, gated by the pre_tool_retrieval config flag (default on). When enabled, the harness searches memory based on the conversation and injects matching notes into context so the model can use prior knowledge without an explicit search. Retrieval does not fire inside sub-agents.

The MemoryBackend trait also exposes init_commands(), recent(limit), search_deep(query), remember(title, content, template), journal(entry), and an optional search_skills_deep(query) for searching skill files alongside memory.

Sub-Agents

Sub-agents are child agent instances spawned by the harness to run a scoped task with a filtered toolset. The Zellij host opens a plugin pane; Montana starts another component instance on a thread in the same process. Both use files to communicate with the parent. Montana peers share its preopened directories and HTTP client; spawning a peer does not create another OS security boundary.

Source: crates/core/src/subagent.rs, crates/core/src/spawn.rs, crates/core/src/subagent_config.rs.

How spawning works

The parent calls the built-in spawn_subagent tool. The harness then:

  1. Writes the task to prompt.md and the filtered tool manifest to tools.json in a new directory under <cowboy_dir>/subagents/.
  2. Requests a peer through the host, passing a JSON-encoded SubagentConfig under the subagent_config key. Zellij loads a visible plugin pane; Montana’s supervisor creates a component instance with its own agent ID.
  3. The child, on its first heartbeat tick, writes a lock file and its pane_id, reads prompt.md + tools.json, and processes the prompt as a user message.
  4. When the child finishes, it writes response.md and removes lock.
  5. The parent detects completion on its heartbeat poll and injects the response back into its own conversation.

The shared spawn protocol uses host-executed shell commands for its spool files. WebAssembly does not itself forbid file I/O: Montana also grants direct WASI access to preopened directories. Pane visibility is a Zellij affordance; Montana has no terminal panes.

Directory layout

Each sub-agent gets a directory named <type>-<short_uuid>:

<cowboy_dir>/subagents/<type>-<id>/
  prompt.md     # task + metadata, written by parent
  tools.json    # filtered tool manifest, written by parent
  lock          # present while running, written by child
  pane_id       # host instance id (a pane id under Zellij)
  response.md   # final output, written by child on completion

<cowboy_dir> is $HOME on a ranch install and $XDG_DATA_HOME/cowboy (default ~/.local/share/cowboy) in lite mode.

Sub-agent types

Three types are defined in the SubAgentType enum. Each restricts the tools the child may call:

TypeAllowed tools
Researchread, search, find, web-search, ls
Coderead, write, search, find, bash, ls, __HASHLINE_READ__, __HASHLINE_EDIT__
Reviewread, search, find, bash, ls

Only Code gets write and the hashline edit tool. Review still has bash, so its tool list does not enforce read-only filesystem access. The configured host and tool policy determine what commands can do.

Models

SubagentConfig::cheap_defaults() selects the child’s provider and model. Its fallback model names are:

ProviderDefault sub-agent model
OpenAIgpt-4o-mini
OpenRouteropenai/gpt-4o-mini
Anthropicclaude-3-5-haiku-latest
Ollamamistral:7b
Codexgpt-5.6-luna
Vercelzai/glm-5.3-flash

An explicit subagent_model wins when that provider has a non-empty key. Otherwise the main harness provider/model is used when it is usable (a key, or a keyless Ollama/Codex provider), falling back to the first keyed provider in the order OpenRouter, Anthropic, OpenAI, Vercel — so the Ollama and Codex defaults above are only reached through the main model.

Status

The parent tracks each child with the SubAgentStatus enum:

  • Starting — directory created, peer starting
  • Running — lock present
  • Completed — response.md present, lock removed
  • Failed(String) — lock removed without a response.md

Polling happens on the heartbeat timer.

Hashline Edit Format

Hashline is the harness’s read/edit format. Each line is tagged with a short hash of its content, so an edit can be rejected when the file has changed since the agent last read it.

Source: crates/core/src/hashline.rs. Exposed as the built-in tools __HASHLINE_READ__ and __HASHLINE_EDIT__.

Format

Lines are formatted as {line_num}:{hash}|{content}:

1:a3|fn main() {
2:7f|    println!("hello");
3:b2|}

__HASHLINE_READ__ returns a file in this form. The agent then references lines by line:hash when editing.

Hash function

The hash is FNV-1a (offset basis 2166136261, prime 16777619) truncated to the low byte (hash & 0xff) and formatted as two lowercase hex characters. With 256 possible values it is meant for change detection, not collision resistance.

Edit operations

__HASHLINE_EDIT__ takes an array of operations. Each references a target line by line:hash:

  • replace — replace lines from start to optional end (inclusive) with new content
  • insert_before — insert content before the referenced line
  • insert_after — insert content after the referenced line
  • delete — delete lines from start to optional end (inclusive)

Every referenced line’s hash is validated against a fresh read of the file before anything is applied. If any hash does not match, the whole edit is rejected and the agent is told to re-read. Operations are applied in reverse line order (so earlier edits don’t shift later line numbers), then the file is rewritten in place with exactly the computed bytes — a file without a trailing newline stays without one, and a blank last line stays. The write itself re-checks that the file still holds the bytes that were read (cksum, in the same shell invocation): if it changed in between, nothing is written and the error names the stale line references. That check needs cksum on the agent’s PATH and allowed by its sandbox policy (Bash(cksum) is in the default allow list); if the check itself cannot run, the edit is refused with the shell’s error rather than a retry hint. Files that are not valid UTF-8 or that contain NUL bytes are refused, since the write cannot reproduce them byte-exact. The parser accepts either operations or edits for the array key, and either content or new_text for the text field.

Browser Automation

The harness can drive a real browser through Camoufox, an anti-detection Firefox fork. Agents navigate pages, read accessibility snapshots, interact with elements, run JavaScript, and take screenshots.

Source: pkgs/camoufox/, modules/camoufox.nix, modules/options/camoufox.nix.

Components

  • camoufox — the Camoufox browser binary (anti-detection Firefox, daijro/camoufox).
  • camofox-server — a Node.js REST server wrapping Camoufox (jo-inc/camofox-browser).
  • cowboy-browser — a shell CLI (pkgs/camoufox/cowboy-browser.sh) that the agent calls; it talks to the server’s REST API.

All three are exposed as flake packages: nix build .#camoufox, .#camofox-server, .#cowboy-browser.

Deployment

The server runs inside a shared NixOS container (cowboy-camoufox) with one camoufox-<user> systemd service per agent. Each agent’s server listens on basePort + (uid - 1338), where basePort defaults to 9377. Browsing happens under a headless Xvfb display.

Enable it with:

services.cowboy.camoufox.enable = true;

This requires services.cowboy.secretsProxy.enable = true — the camoufox container uses the proxy’s veth networking, and the module asserts this.

Options under services.cowboy.camoufox:

OptionDefaultMeaning
enablefalseenable the camoufox container
basePort9377base port; per-agent port derives from UID

CLI

cowboy-browser keeps a session id in a file under $XDG_RUNTIME_DIR and targets the server at $CAMOFOX_URL (default http://10.200.0.1:9377, the veth host address — localhost does not reach the container). Subcommands:

cowboy-browser navigate <url>      Open URL (creates a session if needed), return a snapshot
cowboy-browser snapshot [--full]   Accessibility-tree snapshot of the current page
cowboy-browser click <ref>         Click an element by ref
cowboy-browser type <ref> <text>   Type text into an element
cowboy-browser scroll <direction>  Scroll (up/down/left/right)
cowboy-browser press <key>         Press a key (Enter, Escape, Tab, ...)
cowboy-browser back                Navigate back
cowboy-browser eval <js> | eval -  Run JS in the page ('-' reads from stdin)
cowboy-browser screenshot [file]   Save a screenshot (default /tmp/cowboy-browser-screenshot.png)
cowboy-browser close               Close the browser session

Elements are referenced as @e1, @e2, … from the snapshot output (the @ is optional).

Example

cowboy-browser navigate https://example.com
# snapshot lists interactive elements with @eN refs
cowboy-browser type @e3 "search query"
cowboy-browser click @e7
cowboy-browser screenshot ./result.png

The agent invokes these through its bash tool. cowboy-browser unsets the HTTP proxy variables for its own requests, since the camofox server is local to the veth and does not need credential injection.

API Integration

Provider code formats requests and parses responses in core. The host carries them out asynchronously: the Zellij adapter uses its web-request API, while Montana implements the component’s HTTP import with a native client and sends the completion back as a host event. Provider logic does not depend on a terminal or Zellij panes. Proxy routing, credentials and network policy still depend on the deployment.

Source: crates/core/src/provider/.

Providers

Six providers implement the LlmProvider trait, selected by ProviderType:

  • ClaudeProvider (Anthropic Messages API)
  • OpenAIProvider (Chat Completions / Responses API)
  • CodexProvider (ChatGPT subscription Responses API with SSE)
  • OpenRouterProvider
  • VercelProvider (Vercel AI Gateway, Chat Completions)
  • OllamaProvider (local, keyless)

The LlmProvider trait

The trait does not perform transport I/O. It only serializes requests and deserializes responses:

#![allow(unused)]
fn main() {
pub trait LlmProvider {
    fn name(&self) -> &str;

    fn format_request(
        &self,
        messages: &[Message],
        tools: &[Tool],
        system: &str,
    ) -> (String, BTreeMap<String, String>, Vec<u8>);

    fn parse_response(&self, body: &[u8]) -> Result<LlmResponse, ProviderError>;
    fn parse_error(&self, status: u16, body: &[u8]) -> ProviderError;

    fn set_model(&mut self, model: &str);
    fn api_key(&self) -> &str;
}
}

The request flow is:

  1. format_request() produces (url, headers, body).
  2. Core requests HTTP through its host interface.
  3. The host delivers the correlated HTTP result: a Zellij event or, through Montana, a WIT host event mapped back into core.
  4. parse_response() (on HTTP 200) or parse_error() (otherwise) interprets it.

set_model() supports switching the model at runtime (the /model command).

Credentials

In the managed credential-proxy deployment (services.cowboy.secretsProxy.enable), the agent sends placeholder keys and the proxy injects the real credentials on the wire based on per-domain mappings. Codex is the same pattern: its rotating ChatGPT subscription token is injected outside the agent. Local and ordinary Docker launches hold credentials. See Security Model. For keyless providers (Ollama) api_key() is empty.

Provider defaults

ClaudeProvider targets /v1/messages. Its defaults: model claude-sonnet-4-20250514, max_tokens 8192, extended thinking enabled with a token budget. OpenAIProvider defaults to gpt-4o and auto-selects the Responses API for reasoning models (o-series, gpt-5, codex) to capture reasoning summaries, which are mapped onto thinking content blocks.

Message model

Messages carry a string role and a vector of ContentBlocks, matching the Claude API’s native content-block model. A block’s block_type is one of text, tool_use, tool_result, or thinking.

Errors and retries

ProviderError includes ParseError, ApiError, NetworkError, RateLimited, Timeout, InvalidRequest, AuthenticationError, and Overloaded. is_retryable() marks rate limits (429), network errors, server errors (5xx), and timeouts as retryable; auth and other 4xx errors are not. Automatic retry is wired in: a RetryState with exponential backoff drives a bounded retry of failed LLM calls from the ranch handlers (crates/core/src/handlers.rs, schedule_llm_retry) — a retryable error schedules a backed-off retry until max_attempts is exhausted, after which the failure surfaces as an error message.

Bridges

Bridge services (Discord, email) do not use the provider layer. They reach the agent through the pub/sub message system: inbound messages are polled and processed as user input, and replies flow back out the same way. See Approvals & Outbox.

Plugin Architecture


title: Plugin Architecture tags: [architecture, plugins, extensibility] created: 2026-04-24 updated: 2026-09-06

Plugin Architecture

Cowboy’s NixOS module system is the extension surface. An external module (a “plugin”) adds capability to agents by registering into cowboy’s option tree rather than forking the core. The in-tree Home Assistant skill is the worked example of per-agent content; the examples below use a hypothetical myplugin that runs a service on the host and teaches an agent to drive it.

Everything below is grounded in the implementation under modules/.

The three layers

A plugin separates into layers so that only the top one couples to cowboy:

LayerWhat it isCouples to cowboy?
Infrastructuresystemd services / scripts providing the raw capability (game server, DB, hardware)No — runs standalone
Packagesderivations giving the programmatic interface (Python env, CLI) — built via callPackageNo — self-contained
Harness integrationconditional blocks registering skills/tools/bridges and pushing packages into agent homesYes — guarded by hasCowboy

In myplugin this maps to: services.nix (infrastructure), default.nix (packages, built once and imported by both of the others), and the config block in module.nix guarded by hasCowboy (integration). Disable cowboy and the service still runs.

The plugin contract

1. Own your namespace

Declare options under your own top-level namespace, never inside services.cowboy.*:

# Good
options.services.myplugin = { enable = ...; dataDir = ...; };

# Bad — destabilizes cowboy's option tree and breaks standalone use
options.services.cowboy.myplugin = { ... };

2. Read cowboy state via cowboyLib

Cowboy injects a cowboyLib module argument (modules/lib/default.nix). This is the supported module interface today, not a versioned binary ABI or a promise of compatibility across future revisions. Pin the Cowboy input and evaluate extensions when updating it. Members:

cowboyLib.enabledAgents          # attrset of enabled agents (name -> agentConfig)
cowboyLib.agentNames             # [ "agent" ... ] — for building `agents = [ ... ]` selectors
cowboyLib.forAgents (acfg: {…})  # map an HM config fn over every enabled agent, keyed by user
cowboyLib.forAgentsWhere pred f  # …only agents matching `pred name acfg` (generic haAgents)
cowboyLib.singleAgent "context"  # the one enabled agent, or a lazy throw under multi
cowboyLib.broker                 # human operator user (may be null)
cowboyLib.serviceUser            # cowboy's service-account prefix / shared bridge group
cowboyLib.userFor / homeFor name # agent user / home dir by name
cowboyLib.caBundleFor name       # the trust store the agent's TLS clients must use ($SSL_CERT_FILE)
cowboyLib.inNamespace            # is agent traffic proxied?

Accept it with a fallback so the plugin still evaluates when cowboy is absent:

{ config, lib, pkgs,
  cowboyLib ? { enabledAgents = {}; forAgents = _: {}; },
  ... }:

let hasCowboy = cowboyLib.enabledAgents != {}; in

Per-agent targeting: forAgents and the skills/tools registries fan out to all enabled agents by default. To scope a skill/tool to specific agents, set its agents = [ "name" … ] (empty = all). For per-agent content (prompts that differ by agent), use forAgentsWhere directly — see the Home Assistant pattern.

3. Guard harness integration

Keep infrastructure unconditional; gate cowboy registration on hasCowboy:

config = lib.mkIf cfg.enable {
  # Infrastructure — always
  systemd.services."myplugin-setup" = { ... };

  # Integration — only with cowboy
  services.cowboy.skills = lib.mkIf hasCowboy { ... };

  home-manager.users = lib.mkIf hasCowboy (
    cowboyLib.forAgents (acfg: { home.packages = [ pluginCli ]; })
  );
};

4. Don’t block cowboy.target

Plugin services join the agent lifecycle with a soft pull-in:

  • Use wantedBy = [ "cowboy.target" ] for startup with the target. Add stop and restart coupling only when the plugin needs that lifecycle.
  • Bound restart loops with StartLimitBurst and StartLimitIntervalSec when using Restart = "on-failure".
  • Use requires/after/partOf among the plugin’s own services for internal ordering.

5. Keep packages self-contained

callPackage your own dependencies, and build each one once — a default.nix that both module.nix and services.nix import, rather than two copies of the same derivation kept in sync by hand. Don’t assume cowboy provides any particular package on PATH (the base set is just bash coreutils ripgrep fd jq gh nix git curl, plus camoufox’s browser — see modules/tools.nix).

Extension points

Skills — knowledge + prompt + packages

A skill is a markdown prompt the agent can load, optionally with packages and extra tools. Registered into the global services.cowboy.skills attrset (modules/skills/default.nix):

services.cowboy.skills.myplugin-guide = {
  description = "Drive myplugin from the CLI";
  prompt = ./skills/myplugin-guide.md;  # or promptText = "...inline...";
  requires = [ pluginCli ];             # installed for every targeted agent
  additionalTools = [ "bash" "read" "write" ];
  tags = [ "myplugin" "control" ];
  autoLoad = false;                     # true = loaded at agent startup
  agents = [ ];                         # [] = all enabled agents; or [ "pilot" ]
};

For each agent the skill targets, the module writes ~/.config/cowboy/skills/<name>.md (the harness discovers skills by reading this directory) plus a per-agent skills.json, and installs the skill’s requires packages onto that agent’s PATH — regardless of autoLoad, so an on-demand skill’s binaries are present when the agent loads it. The agents selector filters which agents receive the skill.

Tools — schema’d executable capabilities

A tool is a command template the harness can call with structured args (modules/tools.nix). The default set (bash, web-search, spawn_subagent) merges with anything a plugin adds:

services.cowboy.tools.myplugin-eval = {
  package = pluginCli;
  command = "myplugin run {{script}}";  # placeholders are {{double-brace}}, NOT {single}
  description = "Run a myplugin control script";
  sandbox = "standard";                 # "none" | "standard" | "strict"
  timeout = 60;
  agents = [ ];                         # [] = all enabled agents (defaults always reach every agent)
  schema = {
    type = "object";
    properties.script = { type = "string"; description = "Path to script"; };
    required = [ "script" ];
  };
};

Tools are baked per-agent into the harness WASM and written to a per-agent ~/.config/cowboy/tools.json; the agents selector scopes a tool to specific agents (the default tools — bash, web-search, spawn_subagent — leave it empty, so they reach everyone). A plugin can also expose a capability through a skill alone (prompt + its binary on PATH), which is the lighter path when the agent only needs a CLI.

Plugin registry — be discoverable

Register the plugin in the informational registry so the agent (via its system prompt) and the operator (cowboy plugins / /etc/cowboy/plugins.json) can enumerate what’s installed:

services.cowboy.plugins.myplugin = {
  description = "Host service control";
  version = "1";
  extensionPoints = [ "skills" "units" ];
  units = [ "myplugin-setup.service" "myplugin-daemon.service" ];  # the real units it owns
  agents = [ ];                                                    # [] = surfaced to all agents
};

This is purely descriptive — it wires up no capability, it just makes the plugin enumerable. The per-agent tools.json system_prompt_suffix carries the list to the agent with no harness changes.

Bridges — bidirectional message channels

A bridge connects an external platform to the agent’s pubsub (modules/bridges.nix, options in modules/options/bridges.nix). Each declaration auto-generates a pubsub source plus three systemd units (cowboy-<name>-{ping,ingest,outbox}) under cowboy-bridges.target:

services.cowboy.bridges.discord = {
  pkg = myBridgePkg;                # provides bin/discord-{ping,ingest,outbox}
  env = "/run/agenix/discord-env";  # EnvironmentFile
  icon = "💬";
  approval = {
    required = true;
    notify = "discord";             # which outbox sends the approval prompt
    notify_channel = "1489...";     # channel id within that outbox
    timeout = 3600;                 # auto-reject after N seconds
  };
};

Bridges default to dedicated service accounts. Keep their identities separate from agents and the operator when overriding user. Configure outbound approval on the bridge and inbound routing with routes / defaultAgent. The Security Model explains the separate Redis authorities.

The systemd target

A unit can be started with the agent target:

systemd.services.my-capability.wantedBy = [ "cowboy.target" ];

Per-agent skills: the Home Assistant pattern

The global skills registry distributes to all agents by default and supports an agents selector. For content that differs per agent, the in-tree HA skill (modules/skills/home-assistant.nix) shows the escape hatch: it bypasses the registry and writes the prompt directly into the homes of agents that opted in via a per-agent option (agentOpts.skills.homeAssistant):

let haAgents = lib.filterAttrs (_: a: a.skills.homeAssistant.enable) enabledAgents; in
home-manager.users = lib.mapAttrs' (_: acfg:
  lib.nameValuePair acfg.user {
    home.file.".config/cowboy/skills/home-assistant.md".source = mkHaSkillPrompt acfg;
    home.packages = [ pkgs.curl pkgs.jq ];
  }
) haAgents;

The agents selector handles per-agent selection directly in the registry, so a plain skill that only needs targeting does not need this bypass. HA uses it because its prompt is per-agent content — the endpoint, token, and managed units differ per agent — which a selection-only list can’t express. Use cowboyLib.forAgentsWhere for that case; HA is the worked example.

Worked example: myplugin end to end

myplugin/
├── module.nix      # options + integration layer (skills, agent packages, registry entry)
├── services.nix    # infrastructure: myplugin-setup / myplugin-daemon units
├── default.nix     # packages: the CLI and its client library (callPackage), built once
└── skills/*.md     # skill prompts

Consumer wiring (machines/<host>.nix):

imports = [
  inputs.cowboy.nixosModules.default
  ../modules/myplugin/module.nix
];

services.myplugin = {
  enable = true;
  dataDir = "/home/<user>/myplugin";
  user = "<user>";
};

What lights up because cowboy is present:

  1. The plugin’s skills, written into the agent’s ~/.config/cowboy/skills/.
  2. pluginCli on the agent’s PATH.
  3. The agent’s polkit-managed unit list (managedServices.units) so it can cycle the stack itself.

Managed units must exist. Every name in managedServices.units must be a unit the configuration declares; a polkit rule over a unit that does not exist advertises a control that silently does nothing. modules/services.nix checks this at eval time and warns rather than asserts (services.cowboy.managedServices.validateUnits, default true), because a host’s units may be legitimately conditional — a feature flag, an optional submodule — and a hard assertion would force every consumer to mirror that logic into its unit list. The registry’s plugins.<name>.units list is informational and is not validated.

Checklist for a new plugin

  • Options under your own namespace, with an enable flag.
  • cowboyLib accepted with a { enabledAgents = {}; forAgents = _: {}; } fallback.
  • All cowboy registration guarded by hasCowboy.
  • Packages via callPackage, built once in default.nix; not assumed present.
  • Plugin services declare startup dependencies and bounded restart behavior.
  • A skill’s binaries go in its requires (installed for every targeted agent, on-demand or not).
  • Scope to specific agents with agents = [ … ]; for per-agent content, use forAgentsWhere.
  • Register in services.cowboy.plugins.<name> so the plugin is discoverable.
  • Every unit in managedServices.units actually exists (the validator warns otherwise); plugins.<name>.units is informational.

Different contracts

The Nix module interface configures deployments. The component interface is separately versioned as cowboy:agent@0.2.0; its WIT source defines commands, effects and carriers. The JSON inside a snapshot remains an opaque display model, not a versioned presentation schema. Native control and descriptor records have their own CDDL definitions.

The A2A adapter exposes a limited profile over the root agent. It does not provide general A2A task persistence, peer addressing or every protocol method. See the Component ABI and A2A profile.

Design Overview

Cowboy runs persistent AI agents on infrastructure you control, using Nix to configure their tools, credentials, message access, and approved actions.

The strongest documented deployment is managed NixOS on Linux/x86-64. Local and ordinary Docker launches have different credential and process boundaries. The OCI runtime is a specialized optional route for Cowboy’s own payload. See Deployment Paths for that choice before selecting a host or changing its controls.

Core and the two hosts

The core owns the agent loop, provider request formatting, tools, session state, memory and message polling. Two hosts run it:

HostCore integrationInteraction
Zellij pluginLinks core into a wasm32-wasip1 pluginTerminal input and ANSI display
MontanaLoads a wasm32-wasip2 component through cowboy:agent@0.2.0Commands in and JSON frames out over the native control socket

Zellij implements asynchronous commands, HTTP requests, timers and peer spawning through its plugin API. Montana implements those operations natively and delivers completions to the component. Both use the same semantic intents and provider logic; terminal navigation and pane controls belong to Zellij. Montana runs peers as component instances on separate threads in the same process, sharing its filesystem view and HTTP client.

Two paths to the outside world

Custom asynchronous effects cross the core’s host interface: execution, HTTP, timers and peer lifecycle requests. The component adapter maps these to WIT imports. An effect request does not grant permission to perform it.

Direct filesystem access is a separate path. Montana gives the component WASI preopens for its resolved workspace, home, config, data and temporary directories. Config loading, plan persistence and other direct file operations use those preopens without an execution effect. WASI clocks and randomness are also available; WASI sockets are not enabled.

Montana maps each preopened host directory to the same guest path. This matters because execution effects run native processes: a path computed in the guest must name the same file when passed to a shell command.

To reason about access, inspect both the host’s effect handling and its WASI grants, then the enclosing process identity, mounts and network policy. Sheepdog’s tool-command policy alone does not describe direct guest file access. See Security Model.

Launch branches

cowboy-core
├── Zellij plugin: interactive terminal
└── WASI component → Montana
    ├── managed NixOS: systemd agent service
    ├── local headless launch
    ├── ordinary Docker: image supervisor
    └── optional Cowboy OCI runtime: embedded Montana

The managed daemon runs Montana directly with a per-agent component, under the shared systemd hardening envelope, agent user, configured network namespace, credential proxy and Sheepdog policy. It does not launch the agent through the OCI runtime. Optional socket forwarders use a helper from that runtime package; that does not turn the daemon into an OCI container.

Local and ordinary Docker launches hold provider credentials. The managed proxy deployment keeps them outside the agent. The OCI route has its own mount, seccomp and egress controls; it is not a general container runtime. Tool behavior can differ under these policies. The terms bare, user and container classify a design progression; they are not three switches offered by the runtime binary.

Messages and configuration

Bridge services connect external platforms to Redis streams. Core polls its configured sources and sends replies through their outboxes. Bridge approval can hold an outbound message for a human decision. Recovery does not guarantee completion or preserve reply routing across a restart; see Inbound Message Reliability.

Both hosts pass a flat string configuration map to core. The CLI, persistent JSON, generated NixOS configuration and secret discovery are different inputs; see Configuration. Skills, tools and bridges extend the agent through the module interface.

Implementation map

AreaLocation
Agent state and toolscrates/core/
Terminal hostcrates/harness/
WIT adaptercrates/component/
Headless host and WASI grantscrates/montana/
Specialized OCI runtimecrates/runtime/
Native protocol typescrates/contracts/
CLI and launch selectioncowboy/
Managed services and policymodules/
Credential injectionproxy/
Message bridge frameworkpkgs/bridge/

The Component ABI describes WIT and the native control protocol. The embedding guide explains how to add a host.

Security Model

The strongest documented deployment is managed NixOS on Linux/x86-64. Its agent daemon runs Montana directly under systemd with a per-agent component, an agent user, the shared hardening envelope, a network namespace, the credential proxy and Sheepdog. It does not use the OCI runtime to launch the agent. Local, ordinary Docker and optional OCI launches have different boundaries; see Deployment Paths.

This page explains the controls. The threat model states their conditions and limitations. In particular, Cowboy does not guarantee confidentiality of anything the agent can read. A prompt-injected worker with permitted outbound access can disclose readable data. The agent can also damage its own writable workspace, memory and state.

What is enforced by default

The Nix modules enable the credential proxy by default, and enable Redis ACLs and Sheepdog by default on Linux. These are deployment settings, not properties of an arbitrary core or component build. Managed daemons require the proxy; turning it off is not a supported way to run that daemon with real keys.

Process and filesystem isolation

Each managed daemon runs as its configured agent user. The systemd envelope makes the system read-only, hides other homes, binds the agent’s home back in, and applies private temporary storage, device restrictions, an empty capability bounding set and no-new-privileges. Configured mounts add access to that view. The home, runtime and explicitly writable paths remain mutable.

Inside Montana, direct guest file access uses WASI preopens. These map the workspace, home, config, data and temporary directories at their host paths, so guest reads and native execution agree about filenames. Preopens grant file access separately from custom execution effects; a review of tool policy alone is not a review of all filesystem access.

Montana sub-agents share the host process and its preopens. A filtered child tool list is not a separate OS identity or filesystem boundary.

Network isolation

The managed Linux module creates one shared namespace, cowboy-ns, with a veth pair to the host. The daemons and agent user services join it. It is not one namespace per agent.

Outbound TCP outside the configured subnet is redirected to the proxy. The egress filter allows loopback, the subnet, established replies and DNS to the host resolver, then drops other traffic. Operator-configured SSH destinations and passthrough ports grant direct egress exceptions.

The host Nix daemon is another network authority: its builders run outside the agent namespace. The default denial of direct Nix daemon access is a Sheepdog connect rule. Disabling Sheepdog removes that enforcement. Rebuild requests can instead go through the broker’s configured approval workflow.

Darwin uses a UID-scoped packet-filter redirect and has neither the Linux namespace boundary nor Sheepdog or the Linux Redis ACL implementation.

Credential-injecting proxy

Managed daemons require the proxy and use placeholder provider keys. The proxy reads the real secret from its configured file and injects it on a matching TLS destination route. Exact host routes take priority over wildcard routes; more specific wildcard suffixes take priority over broader ones. A mismatched Host header is rejected before credentials are attached. Upstream certificate verification is enabled by default; the explicit insecure option disables it.

These claims apply to the proxy deployment. Local and ordinary Docker agents hold credentials. Proxy configuration does not remove credentials supplied through some other input.

With egress control enabled, GET, HEAD, OPTIONS and TRACE pass the method gate. Every other method needs an allowed destination or receives HTTP 403. These are method-and-destination restrictions, not a guarantee that external state cannot change. There is no content inspection that prevents readable data from leaving through an allowed request.

Sheepdog policy

On Linux/x86-64, the managed build bakes a per-agent Sheepdog policy into the component’s tool execution path. Sheepdog uses seccomp notification to mediate selected file, execution and connection syscalls. Its policy includes read, edit, create, delete, execution and connection rules. Runtime permission grants can take precedence over baked rules when enabled.

This is conditional on the sandbox being enabled and present in the build. Portable lite builds do not enforce this policy. The optional OCI route uses its own bundle restrictions; it does not inherit all managed-service controls. The policy authority and limitations are specified in Sheepdog policy.

A command can exit successfully even when part of it was denied. Core preserves recognized Sheepdog denial messages in the tool result so the agent can see that a refused read or connection was not an empty successful result.

Per-agent message isolation

With the Linux Redis ACL configuration enabled, each agent has its own inbox and outbox streams and an identity restricted to its name prefix. It cannot read another agent’s streams or edit approval records.

The agent reaches Redis through its own Unix socket at /run/cowboy/<agent>.sock, owned by its user with mode 0600. The host auth proxy selects the Redis identity from the accepting socket and injects the password. The agent holds no Redis password. Per-agent source configuration names these streams and the socket.

Bridges route inbound messages to an agent’s inbox and consume each agent’s outbox. The outbox stream prefix determines the requester; a message’s claimed user does not. Bridge bookkeeping lives outside agent-writable prefixes. Each bridge has its own service user and Redis identity, scoped to its streams and required approval roles. A bridge that creates or resolves approvals still has authority over shared approval records; those roles must be trusted.

This separates agent mail while provider credentials remain shared through the host’s proxy. It is not a tenant boundary between mutually untrusted operators.

Approvals and operational authority

Bridge approvals hold configured outbound messages for a human decision and reject them on timeout. Declared Git effects bind approval to a host-inspected commit and base; apply rechecks policy before pushing. The rebuild bridge has privileged activation authority through its configured wrappers. Treat bridge identities and approval resolvers as part of the trusted host.

See Approvals & Outbox and Gated Effects. After a restart, inspect outcomes before retrying: message recovery can repeat work without restoring the reply’s source binding.

Agent users do not receive journal access by default. Enabling services.cowboy.agents.<name>.allowJournal grants access to host logs as well as the corresponding log-path policy. Logs can contain data from other trust domains. The proxy omits query strings from its own request logs; that does not sanitize output from other services.

System Patterns

Keep credentials behind a service

In the credential-proxy deployment, provider secrets live outside the agent. The agent sends placeholders; the proxy selects the destination route and injects the credential. Local and ordinary Docker launches deliberately hold credentials, so this is a property of the proxy deployment, not of core.

The proxy restricts HTTP by method and destination. With egress control enabled, GET, HEAD, OPTIONS and TRACE pass the method gate; every other method requires an allowed destination. This does not prove that an allowed request leaves external state unchanged.

Cowboy does not guarantee confidentiality of data the agent can read. A prompt-injected worker with permitted outbound access can disclose that data. Keeping provider keys outside the worker does not make its workspace private. See Security Model.

Separate proposal from authority

An agent can prepare a commit or request a rebuild. The bridge and configured host service perform the approved action with their own identity. For declared Git effects, the approval binds the commit and base that the host inspected; the apply step checks them again. See Gated Effects.

Extend configuration before core

Tools, skills, packages and bridges can be registered through Nix modules. Per-agent selectors determine who receives them; the current module interface provides accessors for enabled agents and their homes. Adding such an extension does not require changing the agent state machine. See Plugin Architecture.

Separate shared behavior from host affordances

The two hosts share agent state and semantic commands. Zellij owns terminal navigation; Montana owns the native control socket and component instances. A host must supply both effect handling and the filesystem access its guest uses. Host independence does not imply identical security boundaries.

Component ABI

The agent core (cowboy-core) shares its state machine across hosts. Its custom asynchronous interface is effects out (a Host trait) and inputs in — split into semantic Intents (the shared vocabulary) and raw HarnessEvents (the Zellij transport plus effect completions). This chapter is its ABI: the WIT world cowboy:agent@0.2.0 at specs/wit/cowboy-agent.wit — the single WIT source every bindgen consumer reads — which lets a runtime embed core without linking Rust in-process.

The world is small enough to state in a sentence. It exports one interface, agent, with five functions — initialize, handle-command, handle-host-event, snapshot, describe — and imports one, host, for the effects a component requests: exec, http, set-timer, close-self, spawn-peer, and own-agent-id.

Two host paths

Core has two adapters, and only one goes through WIT:

HostCratePath to coreInputOutput
Zellij plugincowboy-harness (wasm32-wasip1)links cowboy-core natively, calls the Rust traitkeys/mouse/paste → navigation stays local, semantic keys → IntentANSI to stdout (Zellij captures)
Headless embedder montanacowboy-component (wasm32-wasip2)WIT cowboy-agent worldcommand + host-eventJSON display model via snapshot

The Zellij adapter deliberately bypasses the component ABI. It needs keystroke-level fidelity to drive the terminal UI (prompt-line editing, navigation, expand/collapse), and it can link Rust, so paying the component-boundary tax buys it nothing. The WIT world exists for hosts that cannot link Rust — and those hosts don’t remote a TUI, they render the JSON display model and build their own input affordances.

The TUI lives in the adapter. cowboy-core is a headless agent model: it owns the DisplayItem data and AgentHarness::frame_json, while the renderers, syntax highlighting (syntect), modal input, and view state (ViewState) live in crates/harness/src/tui — wrapped around the core as Tui { agent, view }, which derefs to AgentHarness. The generic ANSI atoms (colors, symbols, text, spinner, key types) live in cowboy-ui. Core depends on neither cowboy-ui nor syntect, so the WASI component sheds both. Every view↔core coupling crosses the Host seam as a defaulted hook (below), never a field poke.

The named-intent seam

Raw keystrokes never cross the ABI. Core’s input surface splits in two:

  • Navigation / view — scroll, expand/collapse, search, prompt-line editing. Client-local: the Zellij adapter owns it (it holds the TUI projection); a JSON client navigates its own copy of the frame. Never crosses WIT.
  • Semantic intents — the handful of actions that change agent/session state: submit / interrupt / approval / debug / quit. The type is Intent, owned by cowboy-contracts and re-exported by cowboy_core::intent so the native adapters and the component adapter cannot drift.

Both adapters converge on one vocabulary via AgentHarness::handle_intent — no bifurcation. The Zellij adapter’s semantic key arms (in crates/harness/src/tui/input: prompt.rs Enter → Submit, navigate.rs Esc → Interrupt and d → Debug, mod.rs sentry y/o/A/n → Approval) call the same method the WIT component maps command’s cases onto. Intent derives serde, so it is also the montana driver’s inbound wire shape ({"submit":"hi"}, "interrupt", {"approval":"once"}, "debug", "quit", specified by specs/control.cddl) — one vocabulary end to end, not a second protocol.

Commands vs. host events

The world’s inbound surface is two types, delivered through two exports, because they have different producers and different trust semantics — a transport command must never be confusable with an effect completion:

  • command — semantic operations an operator or protocol adapter applies to the instance: submit(string), interrupt, approval(approval), shutdown. Delivered by handle-command.
  • host-event — completions and notifications the host produces: timer, command-result, web-result, permission-result. Delivered by handle-host-event.

Both return bool: true when observable state changed and a fresh snapshot is worth pulling.

debug is deliberately not a WIT command — it opens a local log viewer, so it is a host diagnostic rather than an agent operation, and it stays in the native Intent vocabulary only. Intent::Quit is what command.shutdown maps onto guest-side.

Neither type carries key, mouse, or paste. Core keeps the full Key/KeyEvent/Pasted + HarnessEvent types for the Zellij transport.

One component instance has one foreground turn. The world exposes no task handle, no detach/background operation, and no reattachment; submitting while work is active queues the input, and interrupt asks the guest to abandon its turn without promising cancellation of already-issued host effects. Peers spawned via host.spawn-peer are separate instances, not background tasks of the spawner.

Rust-only host verbs

Some seams stay in the Rust Host trait and never enter the world at all, because they are meaningless off a terminal.

The pane verb is one: pane visibility is client-local view state a windowed host acts on, so montana/remote never needs it. pane_op lives in the Rust trait (the native Zellij adapter implements it); the WASI component no-ops it.

Defaulted view hooks

Four more Rust-only Host methods carry the view couplings — all defaulted, so headless hosts inherit them for free (they carry no meaning off a terminal), and only ZellijHost overrides them:

  • reveal_latest() — “a new item was appended; snap the view to it.” Core calls it wherever the view must follow appended output; ZellijHost records a bit its Tui drains after the event (honoring the user’s mode).
  • open_debug(path) — ZellijHost opens the debug log in a floating pane; headless no-ops. Keeps the zellij action new-pane … less +F exec out of core.
  • has_draft_input() -> bool — core’s idle judge/autoreprompt guard; the draft lives in the adapter’s view state, so the default is false.
  • editor_active() -> bool — core’s heartbeat cadence; the $EDITOR round-trip is a Zellij-only affordance, so the default is false.

ZellijHost and its Tui share an Rc<HostSignals> cell block: core→view for reveal_latest, view→core for has_draft_input/editor_active. Like pane_op, none of these are in the WIT world — they are Rust-trait seams a windowed host implements and everyone else ignores.

The opaque-correlation rule

Every effect that expects a reply carries a correlation — an opaque list<tuple<string, string>> created by the guest (core holds it as Ctx, an ordered map, and converts at the boundary). The host echoes it back verbatim on the matching host-event and never reads or writes it. The request “kind” lives inside it as a string, meaningful only to core (see dispatch.rs).

A correlation identifies an effect callback, not a foreground or background agent task — it is not a handle a host can hold, poll, or reattach to.

The payoff: new command/web kinds are new correlation string values, not new ABI surface. The world never changes when core grows a new internal effect kind. A host that treats the correlation as bytes-in/bytes-out is forward-compatible by construction.

Effects are requests, not grants

An exec or http call across the host import is a request. Nothing about crossing the boundary authorizes it: the host evaluates it under the active backend and capability policy (sheepdog, the credential proxy, the OCI sandbox — see Security) and may refuse. The same is true of initialize’s config map: its entries are configuration, not capability grants. own-agent-id is the one purely informational import — the instance’s own id, which core’s heartbeat uses to address itself.

Additive timer semantics

host.set-timer(secs) is additive and one-shot: each call arms an independent timer, and there is no cancellation. This mirrors Zellij’s set_timeout, which has no cancel API. Core’s heartbeat (heartbeat.rs) compensates with an armed_count that culls the resulting fan-out down to a single live chain.

A host must not “improve” this to one cancelable timer. Core’s re-arm logic assumes additive delivery; a cancel-on-rearm host would stall or double-fire the heartbeat. A headless embedder implements this as a plain deadline min-heap: push on set-timer, pop-and-deliver when the earliest deadline passes.

The JSON frame contract

snapshot() returns one line of literal JSON:

{"status": "WaitingForInput", "items": [ /* DisplayItem… */ ]}

status is AgentStatus; items is the DisplayItem list core also renders to ANSI. ANSI highlight caches (highlighted_lines, highlighted_input) are #[serde(skip)]’d — they are render artifacts, not state.

Deliberately unstabilized. WIT specifies only the carrier — a string holding one JSON object. The shape inside tracks core’s display model and will change as that model settles; a typed, versioned presentation contract is deferred work, so no host may treat this serialization as a stable interface by accident. montana’s ndjson stream wraps these frames per agent as {"seq", "agent", "frame"} — that envelope is specified, by the control-frame rule in specs/control.cddl.

The describe contract

describe() returns the protocol-neutral agent descriptor, also as one JSON object in a string: name, description, version, provider, model, capabilities, and the skill list (one entry per tool, carrying its JSON Schema input schema verbatim). Unlike the snapshot, this record is specified — by specs/descriptor.cddl, with Descriptor in cowboy-contracts as the Rust shape.

It exists so a host can advertise the agent without knowing anything about Cowboy’s internals: crates/a2a maps this one export onto an A2A agent card and its skill list, and montana serves that mapping. Protocol adapters are bindings onto describe + handle-command, not alternative agent APIs.

/connect — a harness as a client of another harness

/connect <host:port>#<secret-path> (or /connect <montana-uds-path>) turns a Zellij harness into a thin client of a remote agent, built entirely on the two contracts above: inbound, montana’s ndjson envelope rehydrates into Vec<DisplayItem> (pinned by display_item_round_trips_through_json); outbound, every user action is one Intent wire line (pinned by intent_wire_shapes). The fragment names a bearer-token file — the token is the first line of the connection, exactly the cowboy-runtime __forward handshake; the token itself never enters plugin config or argv.

Mechanically it is drop-and-connect: the plugin writes a request JSON to connect_request_path and quits; the Python launcher’s relaunch loop (cli._run) starts a fresh Zellij whose plugin sees connect_target at load() and comes up in client mode (crates/core/src/client.rs) — no providers, no session, no API keys. Because a WASM plugin cannot hold a socket, a cowboy _connect bridge process owns the connection and streams each frame line in via zellij pipe --name connect-frame; sends go out as one-shot cowboy _connect --send connections. /disconnect is the same sentinel dance back to a local session. Standalone entry: cowboy connect TARGET#SECRET-PATH.

Caveats: in a full NixOS (netns) deployment, outbound TCP is DNAT’d to the secrets proxy — a /connect target port must be in secretsProxy.passthroughPorts (UDS targets are unaffected; sheepdog always allows AF_UNIX), and the in-session handoff loop exists only in lite launches (full mode: run cowboy connect from a host shell — a client session needs no agent config anyway). A read-only __forward --ro port works as an observer: frames flow, intents are dropped server-side.

Known Zellij leaks in core

A few Zellij-isms survive in cowboy-core. None break a headless host; they degrade gracefully and are tracked here for eventual promotion to proper host verbs. (The two obvious candidates are already out: the $EDITOR round-trip is Tui-local, and the debug pane is the open_debug host hook.)

  • Subagent wasm path (spawn.rs) — the fallback peer path assumes a Zellij plugin layout (data_dir()/zellij/plugins/cowboy-harness.wasm). A headless host overrides it via config / AGENT_HARNESS_WASM.
  • Model-facing prose (config/loaders.rs) — the spawn_subagent tool description says “in a separate Zellij pane”. Cosmetic; visible to the model, not load-bearing.

Filesystem access alongside WIT

Core also uses direct filesystem operations for configuration, plan state and metadata. Montana grants these through WASI preopens, mapping host directories to identical guest paths so native execution effects see the same files. These grants are separate from the custom WIT imports. An embedder must review both the preopens and effect policy; see Design Overview.

Hacking

Cowboy’s core is a headless state machine shared by the Zellij plugin and Montana. Hosts supply asynchronous effects and the filesystem access the guest uses. The optional OCI runtime embeds Montana. This guide explains where to extend that implementation.

You want toStart from
Run the agent under your own UI, transport or runtimeThe WIT world and templates/host
Ship the agent to a machine without linking Rustpackages.agent-component
See a complete host with threads, sockets and peerscrates/montana/
Add an effect the agent can performThe Host trait in crates/core/src/host.rs

The contract: the WIT world

specs/wit/cowboy-agent.wit defines cowboy:agent@0.2.0, the one interface every non-Zellij host programs against. It exports agent — initialize, handle-command, handle-host-event, snapshot, describe — and imports host — exec, http, set-timer, close-self, spawn-peer, own-agent-id. Commands and host events are separate exports so a transport message can never be mistaken for an effect completion; effect results carry an opaque correlation the host echoes back unchanged; and an effect request is not an authorization, so the host stays responsible for policy. The Component ABI chapter walks through each of those rules; the file itself is short enough to read first.

WIT defines the custom interface; the guest also needs WASI filesystem grants. The JSON inside snapshot is an opaque display model, not a stable schema. The config map given to initialize is the flat string map core parses; crates/contracts/src/keys.rs inventories recognized keys. A new host-provided capability needs an import and implementations in the affected adapters.

The artifact: packages.agent-component

Build from source; see Installation for release availability. nix build .#agent-component produces share/cowboy/cowboy-agent.wasm, the agent core compiled to wasm32-wasip2 and wrapped by crates/component/ so it exports exactly the world above. The wrapper, crates/component/src/lib.rs, is small: it maps initialize onto AgentHarness::on_load, handle-command onto AgentHarness::handle_intent, handle-host-event onto AgentHarness::handle_event, and snapshot and describe onto AgentHarness::frame_json and AgentHarness::describe_json, and it implements the core’s Host trait by calling the world’s imports. A second output, agent-component-lite, is the portable build without the NixOS-only pieces.

An external flake gets it as cowboy.packages.<system>.agent-component. That is how a host stays on one cowboy revision for both the component and the WIT it was built from.

The starting point: templates/host

nix flake init -t github:dmadisetti/cowboy#host
nix develop --command cargo run

The template is the smallest embedder that runs: a flake taking cowboy as an input, and one Rust binary (templates/host/src/main.rs) that instantiates the component with wasmtime, answers the six imports in-process, calls initialize with a tiny config map, submits one prompt, prints every snapshot as a JSON line and exits. Its HTTP handler returns an error, timers are delivered without waiting, and peer spawning is ignored. Replace those stubs before using it for real agent work. Its devShell links wit/ to the cowboy input’s specs/wit so the bindgen! bindings track the same revision as the component. templates/host/README.md explains each export, each import, and what to replace to put a real UI, HTTP client and clock around the loop.

The headless reference: montana

crates/montana/ is the host cowboy itself ships: the daemon under systemd on NixOS, the process inside the container image, and the library the OCI runtime (crates/runtime/) calls in-process. crates/montana/src/agent.rs implements the imports with threads — exec and http run off the guest thread and post their results back into one channel, timers are additive sleepers — and crates/montana/src/lib.rs assembles the config map, resolves and preopens the agent’s directories at their own paths, and starts the control socket. Frames stream out as ndjson and intents come in the same way (specs/CONTROL.md); the optional A2A front in crates/a2a/ is a protocol adapter over describe and handle-command (specs/A2A.md). Read it when the template’s answer to a question is “a real host does more here”.

Inside the component: the Host seam

Hosts that can link Rust do not need the component at all. The Zellij plugin (crates/harness/) links cowboy-core directly and supplies a Host implementation of its own. That trait, in crates/core/src/host.rs, is the seam the WIT world mirrors. Commands and HTTP requests return results later as events correlated by an opaque Ctx map. Timers, pane operations and peer lifecycle requests also go through the host. AgentHarness::with_host takes the implementation; the entry point after that is AgentHarness::on_load, in crates/core/src/harness.rs.

For a new asynchronous effect, update the Rust host trait and its adapters. If component hosts need it, update WIT and the component and Montana bindings as well. Terminal-only view hooks can remain Rust-side; headless adapters have no terminal view to update.

What is not an extension point

  • Configuration keys are parsed in one place, crates/core/src/config/; a new knob is a new key there and a new line in the launcher that emits it, not a host-specific side channel.
  • Tools and skills have configuration and module interfaces; see Plugin Architecture.
  • A host’s boundary includes effect policy, WASI preopens and the enclosing process permissions, mounts and network controls. The Security Model describes the managed deployment; a new embedder must supply its own controls.

Concurrency Locking

The agent runs one session as an event loop. SessionLock (crates/core/src/lock.rs) keeps LLM requests, tool execution and tool-output summarization from interleaving, and queues user input that arrives while any of them is in flight. Every hold on the lock is bounded: one that goes too long without evidence of progress is reclaimed by a sweep the heartbeat runs.

What holds the lock

SessionLock tracks three kinds of hold:

  • llm: Option<Lease> — an LLM request is in flight.
  • summarizing: Option<Lease> — a tool output is at the summary model.
  • active_tools: HashMap<String, ActiveToolState> — tool executions keyed by their unique call_id, each with its own started_at and timeout.

is_locked() is true while any of them is set. at_tool_boundary() is the inverse — no request, no summarization, no registered tool — and is the only point at which queued input is replayed.

Leases and epochs

A Lease is a hold with a deadline: started_at, a timeout, and an epoch naming this particular hold. The budgets live in lock::lease: LLM (300s) covers one round trip with no intermediate signal, SUMMARIZING (180s) one cheap-model summarization of a single tool output.

The deadline means “time since the work last moved”, not “time since dispatch”. touch_llm() resets started_at on evidence the turn advanced — a stream poll whose transcript grew, a compaction reply that starts its second round trip — so a long but healthy streamed turn is never reaped, while a relay that answers every poll with an unchanged buffer does not push the deadline out.

try_acquire_llm(timeout) takes the LLM lease and returns its epoch, or None when a request is already out. acquire_or_current_llm(timeout) returns the live epoch when the lease is already held and takes it otherwise; compaction uses it because it is reached both from inside a turn and from a free lock. set_summarizing(timeout) takes the summarization lease and returns its epoch. force_release_llm() and clear_summarizing() release them.

Epochs come from one monotonic counter shared by both leases, so an epoch names exactly one hold for the life of the session. A dispatch stamps its epoch into the request’s Ctx (dispatch::stamp_epoch, key EPOCH_KEY) and the host echoes it back with the reply. Before any reply reaches a handler, AgentHarness reads it (dispatch::ctx_epoch) and asks epoch_is_current(epoch); a reply whose lease has been reaped or re-taken belongs to a turn that has moved on and is dropped. A reply carrying no epoch passes through — failing closed there would turn a host that drops the key into a starved agent.

Tool tracking

register_tool(call_id, tool_name, timeout) records an ActiveToolState; complete_tool(call_id) removes and returns it. Results are matched by call_id, never by tool name, so concurrent calls to the same tool are tracked independently. set_exec_kind(call_id, kind) attaches the execution-policy classification made at dispatch, so a continuation of the same call is held to the same policy. release_all_tools() drops every registration at once and hands the states back: an interrupt settles the whole set, and the caller still owes each one an abandoned-process record.

The sweep

sweep_expired() reclaims every hold whose deadline has passed — the LLM lease, the summarization lease and timed-out tools alike — and returns them as ExpiredLock values (Llm { waited, epoch }, Summarizing { waited, epoch }, Tool(ActiveToolState)). The leases come first so their handlers settle the turn before a tool result handled afterwards can start the next request.

AgentHarness::sweep_session_lock (crates/core/src/heartbeat.rs) runs it at the head of every timer fire, before any branch can return, and settles each entry:

  • an expired LLM lease drops any stranded relay stream, resets an active status to WaitingForInput and drains queued input; a late reply is then dropped by the epoch guard;
  • an expired summarization routes a synthetic 504 through handle_summarization_response, which truncates the output it still holds, appends the tool result the model is waiting on and continues the turn;
  • a timed-out tool gets a synthetic tool result saying the harness stopped waiting, naming the PID that was not cancelled — Host::exec is fire-and-forget, so a timeout can only stop waiting — and the process is tracked as abandoned.

hold_summary() renders what holds the lock and how long each hold has left; submit logs it whenever input is queued, so a quiet agent says what it is waiting on.

Input queueing and coalescing

submit (crates/core/src/intent.rs) calls queue_user_input(content) when input_can_run_now() is false. boundary_reached() returns None unless at_tool_boundary() holds, and otherwise drains the queue through drain_ready_inputs():

  • a single queued input is returned verbatim;
  • several are coalesced into one message under a [Multiple inputs received while processing:] header, joined by ---.

execute_next_tool_call calls boundary_reached() after a tool result is recorded, gated on the same input_can_run_now() predicate submit queues on, so a rate-limit backoff — which holds no lease — cannot let queued input through. The heartbeat’s tick also drains the queue (check_queued_inputs, behind the same predicate), so input stranded by a lock-free window is picked up without another turn event.

Tests

lock.rs covers LLM mutual exclusion, boundary detection, coalescing, tool timeout removal, the combined locked-state predicate, the sweep reclaiming all three kinds of hold in lease-first order, touch_llm keeping a live stream from being reaped, and an epoch ceasing to be current once its lease is gone.

Result Summarization

Large tool outputs are stored on disk as recallable artifacts and replaced in the context window with a short stub. When summarization is enabled, the stub’s body is produced by a cheap “summary” model instead of blunt truncation.

Implementation: crates/core/src/tools/summarize.rs (classification, prompts, request formatting, stubs, fallback truncation) integrated through harness.rs (handle_command_result) and handlers.rs (handle_summarization_response).

Threshold

A tool output is summarized when its length exceeds DEFAULT_SUMMARIZE_THRESHOLD = 12000 characters; no launch-configuration key changes that. Input sent to the summary model is capped at MAX_SUMMARIZABLE_LENGTH = 50000 characters.

Summary models

The summary model’s provider is taken from the summary_model config and resolved through the shared key chain. Default model per provider:

ProviderDefault model
Anthropicclaude-sonnet-4-20250514
OpenAIgpt-4.1
OpenRouteropenai/gpt-4.1

Those three are the only summarizers there are (ResolvedProvider); there is no Ollama or Codex summarizer. ollama — like any unrecognized provider name — resolves to the OpenAI key and the OpenAI request shape, so a summary_model of ollama:<model> does not reach a local Ollama server, it sends <model> to OpenAI. codex resolves to the first key available in the order OpenAI, Anthropic, OpenRouter, and because that is not a direct hit the configured model name is dropped in favour of that provider’s default above. With no key at all, no summarizer is built and outputs are truncated instead.

Requests use temperature: 0.3 and max_tokens: 2048 (or the provider-appropriate completion-tokens parameter for OpenAI-compatible APIs).

Output classification

The tool name determines an output type, which selects a type-specific prompt and fallback head/tail ratio (ToolOutputType::from_tool_name):

TypeTool namesHead ratio
FileContentread, cat, __HASHLINE_READ__0.7
SearchResultsgrep, rg, search, ast-grep0.5
DirectoryListingls, find, fd, glob0.3
CommandOutputbash, sh (and any unknown tool)0.6
StructuredDatanix-search, gh0.5
WebContentweb-search, web-fetch0.5

Each type has its own prompt: command output preserves errors and exit codes verbatim, search results are grouped by file with line numbers, directory listings are grouped by kind, structured data extracts names and versions, and web content extracts main facts and quotes.

Code-file handling

For FileContent, SummarizationRequest::code_file splits the file into three parts: a before-section and after-section (summarized) around a relevant range that is reproduced exactly with hashlines. When no range is supplied the relevant range defaults to the middle third of the file. The prompt instructs the model to keep the surrounding summaries to one or two sentences and preserve the hashline section verbatim.

Pipeline

  1. A tool result arrives in handle_command_result. The full output is logged to the session’s messages.jsonl exactly once, regardless of what enters context.
  2. If the output exceeds the threshold and the session is ready, it is stored as a disk artifact and kept in memory for UI expansion.
  3. If summarization is enabled, a PendingSummary is queued, the request is dispatched to the summary model, and set_summarizing(true) gates user input until the response returns.
  4. In handle_summarization_response, the lock is cleared, the provider-specific response is parsed, and the summary is wrapped in an artifact stub. On a parse failure or non-200 status, fallback_truncate is used instead.

If summarization is disabled, step 3 is skipped and the artifact stub is built directly from fallback_truncate.

Artifact stubs and recall

format_artifact_stub wraps the summary with a header and a footer noting that the full output is stored:

[Artifact <call_id> | <tool> | <bytes> bytes, <lines> lines]
<summary>
[Full output stored — use recall_artifact("<call_id>") to search or read more]

The agent retrieves the full output on demand through the recall_artifact builtin tool (wire name __RECALL_ARTIFACT__), which reads or searches the stored artifact by call_id.

Fallback truncation

fallback_truncate splits at line boundaries using the type’s head/tail ratio, operates on character counts to stay UTF-8 safe, and inserts a [...N chars omitted...] marker between the kept head and tail. It is used whenever the summary model is disabled, unreachable, or returns an unparseable response.

Inbound Message Reliability

A sender can receive no reply even though the agent started the work. A restart can also cause work or an external action to happen again. Before resubmitting a silent request, the operator must inspect the session, outbox, bridge outcome and any action the request could have caused.

Redis recovery retries delivery to dispatch, not to successful completion. The agent queues acknowledgment when it hands a message to the agent loop, without waiting for an answer. Once that acknowledgment lands, Redis will not recover an interrupted turn. If it does not land, reclaim can dispatch the message again. There is no exactly-once guarantee for completion or effects.

Session recovery is separate. On restart, the agent resumes its saved session, and an idle reprompt can act on the unanswered request again. The active message and reply-source binding lived only in process memory. The work can recur while the reply loses its source binding: a resumed answer may reach the session but no source outbox, leaving the original sender with silence.

ObservationWhat it establishes
Service is runningThe process is up; useful progress still needs inspection
No pending Redis entriesNo entries await acknowledgment in that group
Acknowledgment landedDispatch was acknowledged, not completion
Reply in the agent outboxA reply was queued for the bridge
Session resumedConversation state returned; source routing may not have

There is no durable deduplication across agent restarts and no durable binding that reconnects a resumed turn to its source. Within a process, a bounded seen window suppresses repeat dispatch. Failed acknowledgments retry the acknowledgment first; exhausting that budget allows reclaim to offer the request again. External effects are not transactional with acknowledgment.

The operator procedure

Use the host-side operator identity to inspect missing or duplicate replies. Check the agent with cowboy status and cowboy doctor first. A running service does not prove that a request completed or reached its sender.

The identity. Not the agent’s. Every agent credential is scoped to ~<agent>:* and holds only the verbs the poll issues, which does not include XPENDING — the agent has no use for it, and an agent that could enumerate its own pending list still could not see another’s. modules/pubsub.nix generates one identity for a person instead: operator, readable across the whole keyspace, with no @admin (it cannot rewrite the ACL) and no @dangerous (no CONFIG, FLUSHALL, KEYS, SHUTDOWN). Its URL is root-only, and Redis binds inside the namespace, so every command below is ip netns exec:

sudo -i
. /run/cowboy/operator-redis.env          # sets REDIS_URL
r() { ip netns exec cowboy-ns redis-cli -u "$REDIS_URL" "$@"; }

Three names come from the generated configuration rather than from memory. For agent <agent> and source <source>, /etc/cowboy/<agent>/sources.json names the stream; the group is cowboy; the consumer is harness-<source> — the SOURCE, not the agent (RedisStreamsSource::new names it after the source it polls; two agents on the same source carry the same consumer name on their own separate streams):

stream=$(jq -r '.sources["<source>"].stream' /etc/cowboy/<agent>/sources.json)
group=cowboy
consumer=harness-<source>

1. Is anything stuck? The pending list is the only record of an entry the agent read and did not acknowledge:

r XPENDING "$stream" "$group"

A summary of 0 means the group has no pending entries. It does not prove that every stream entry was read or that dispatched requests were answered.

2. Which entry, and how long has it been stuck? The extended form lists one line per entry — id, consumer, idle milliseconds, delivery count:

r XPENDING "$stream" "$group" - + 10

An entry whose idle time is under ten minutes is not yet eligible for reclaim. It may be queued, or its reader may have stopped; age alone does not prove progress. Reclaim waits until RECLAIM_IDLE_MS has passed. An entry idle for longer than that whose delivery count is not rising is one no poll is reaching — check that the agent is running at all before doing anything else.

3. What was in it?

r XRANGE "$stream" <id> <id>

4. Did the turn already have an effect? Ask this before step 5, because step 5 is the one action in this procedure that can do harm. The agent’s reply goes to its own outbox, and that stream is the record of it:

r XRANGE "<agent>:<source>:outbox" - + COUNT 20

An outbox entry shows a reply was queued, not that the external platform delivered it. Inspect bridge delivery state and the destination before resubmitting. Also check files, rebuild results and other requested actions.

Do not assume reclaim will discard a request that already had an effect. Deduplication is bounded and lives in process memory; after restart the same entry can dispatch again. Neither reclaim nor manual resubmission is an exactly-once operation. If work already happened, reconcile that outcome rather than blindly replaying the original request.

5. Resubmit only after reconciliation. If retrying is appropriate, submit a new entry. Keep in mind that the old pending entry can still be reclaimed. This minimal command submits content only; it does not restore the original sender or channel metadata and therefore does not promise a routed reply:

r XADD "$stream" '*' content '<the content from step 3>'

Implementation

What one poll does

poll_command emits three redis-cli calls in one sh -c:

  1. XGROUP CREATE … 0-0 MKSTREAM, idempotent, so a stream that does not exist yet is not an error.
  2. XAUTOCLAIM <stream> <group> <consumer> <idle> <cursor> COUNT 10 — the recovery step. XREADGROUP > moves an entry into the group’s Pending Entries List under the consumer that read it, and never offers a pending entry again — > means “entries no consumer has read”. So an entry a crashed run had taken is invisible to every subsequent >, and the reclaim is the only thing that hands it back. The consumer name is deliberately stable across restarts (RedisStreamsSource defaults it to harness-<source>, and the comment there says why): a fresh name per process would leave the whole previous pending list keyed to a consumer that will never ask for anything again. Stability keeps the list reachable; the reclaim is what reaches it.
  3. XREADGROUP GROUP … COUNT 10 STREAMS <stream> '>' — new entries.

Reclaim before read, so a message stranded by a restart is not queued behind messages that just arrived. A marker line separates the two replies; both are parsed by the same hand-rolled reader over redis-cli --no-raw output.

The scan carries its cursor. XAUTOCLAIM scans the pending list from a cursor and answers with the one to resume from. A scan that spends its COUNT on entries too young to claim returns an empty batch and a non-zero cursor; restarting from 0-0 every poll re-walks that same head, and an orphan behind a full scan’s worth of young entries is only reached once the head drains on its own. RedisStreamsSource::reclaim_cursor keeps the returned cursor, which is why poll_command and parse_poll_result take &mut self. Redis answers 0-0 when the scan wraps, so the walk returns to the head on its own.

An entry with no usable content is still emitted, with an empty one. It has to be: an entry nothing returns is an entry nothing acknowledges, and since the poll reclaims the pending list it would come back on every poll forever, masking the ten entries behind it. The manager acknowledges and drops it.

Acknowledgment semantics

XACK answers with the number of entries it removed from the pending list. The reply is examined, not the exit code alone: redis-cli prints a refused command’s error text on stdout and exits zero, so an ACL that does not grant XACK used to read as an acknowledgment that landed. The generated command requires the reply to be a run of digits — zero included, which is the idempotent case — and exits ACK_REPLY_NOT_A_COUNT otherwise.

The acknowledgment is queued at dispatch, not at completion. process_pubsub_messages calls SourceManager::mark_processed on each message as soon as the prompt has been handed to process_user_input, and the same poll tick then calls send_pubsub_acks. Nothing waits for the turn to produce an answer. Redis’s redelivery therefore covers exactly one window — from the XREADGROUP that read the entry to the XACK that follows its dispatch — and nothing after it. A process that dies mid-turn has already acknowledged the entry it was working on; the pending list holds no record of it, and no reclaim will bring it back. That is the deliberate trade: the alternative, acking at completion, makes every crashed turn a guaranteed re-answer, including the ones that had already sent their reply.

A failed acknowledgment retries the acknowledgment, not the request. SourceManager::retry_ack queues the same XACK again, up to MAX_ACK_ATTEMPTS. Only once that budget is spent does the manager forget the entry, which hands it back to the reclaim. Forgetting on the first failure — before the retry budget is spent — gives the reclaim a message the agent has already answered, and the next poll starts a second turn on it.