Troubleshooting
Start with cowboy doctor. A service can be up while the agent has stopped
reading messages.
cowboy list
cowboy status
cowboy doctor
cowboy doctor dev --json
The list is the registered inventory. Status observes running instances and last reported activity. Doctor separates three observations:
| Observation | What it establishes | What it does not establish |
|---|---|---|
| Process/service exists | The host still has a running process | The event loop is making progress |
| Event loop responds | A diagnostic request produced a new frame | Inbox polling still runs |
| Inbox consumer is active | Redis records recent reads | The model finished a turn or a reply reached its destination |
A replayed status frame can outlive a stalled loop. Doctor asks for a new one. Its attention check examines inbox consumer activity independently.
Checks report ok, fail, or unknown. Unknown means the probe could not establish an answer, for example because the caller cannot read an agent’s Redis socket or the deployment has no supported probe. Read the reason; do not translate it into healthy. On a managed host, rerun as an authorized operator with the required privilege if the reason is access. Do not broaden socket permissions to make a diagnostic green.
Find the right log
# Managed service:
systemctl status cowboy-serve-dev --no-pager
journalctl -u cowboy-serve-dev --since today
systemctl list-units 'cowboy-*'
# Backgrounded local serve (default headless name):
cowboy logs agent
# Docker walkthrough:
docker logs cowboy-eval
Inspect the first failed service and its dependencies. The managed daemon requires the namespace and credential proxy; it does not launch an OCI bundle. A cage startup failure instead belongs to Docker, its runtime registration, the seed/rootfs, bundle policy, or companion services.
The CLI or Zellij is missing
Use the source-build installation. Select the CLI package explicitly; the default flake package is not the Python launcher. A Python source install alone does not bundle the harness WASM. Zellij is required for interactive sessions and the connect client, not for Montana’s headless service.
The client opens but the model fails
Check the provider/model specification, the credential source for this route, and the upstream response. Local and ordinary Docker paths hold real keys; managed agents use placeholders and need the proxy’s secret path to be readable by its service user. User and system keys files accessible to group or others are ignored. See secret discovery.
On a managed daemon, a runtime provider switch also needs provider access in
the Nix configuration. A proxy credential alone does not supply the component’s
provider key placeholder. Declare the provider through a model option or
extraProviders and rebuild.
Record the provider, model, HTTP status, and request ID without copying keys. A method or destination denial calls for reviewing the configured policy, not bypassing the proxy.
A tool is denied
Identify the layer from its result: tool approval, filesystem permissions, Sheepdog syscall policy, service hardening, proxy methods/destinations, or OCI policy on the cage route. Compare the operation to that deployment’s declared access. A successful local run is not evidence that the managed policy should permit it.
An approval or reply seems stuck
Follow the request through ingest, the agent inbox, the agent turn, and the outbox. For Discord, the receipt reaction establishes that the bridge queued the message, not that the agent answered. Check doctor for continued inbox consumption and inspect whether an approval is pending, rejected, expired, or superseded. Only configured approvers can resolve a gated Discord request.
For Git effects, preserve the request ID and SHA. The approval must match the runner’s description and the applied SHA. A branch moving after review causes a refusal; fetch and prepare a fresh request instead of treating the old approval as permission for another commit. A push can succeed while its verification fetch fails: inspect the remote before retrying.
After interrupted work, a request may have no reply, be redelivered, or resume without its original reply destination. Inspect external results before asking for the same side effect again. See message reliability.
A rebuild succeeds but nothing changes
The rebuild service compares the built system closure with the running system, or the new home activation path with the prior one when it can read that state. An identical result gets an explicit no-change note. A successful command can therefore mean no new deployment.
First check the rebuild type: a home rebuild cannot deploy system services, networking, or system packages. Check that the approved commit contains the change, that it reached the tracked branch, and that the flake imports the edited file. Check resolved input overrides in the approval and result. Then check the deployed generation and whether the daemon restarted onto the new component. Doctor compares the generated agent configuration timestamp with the daemon start time to flag possible configuration drift. If the rebuild could not establish a prior home generation, absence of a no-change note is not proof of change.
For settings-based containers, an existing result link is reused at startup.
Build the edited generated configuration with cowboy bake inside the container
before restarting; merely editing JSON or replacing the base image does not
necessarily replace that generation.
Restart without losing the evidence
Copy mutable state before repair. Managed daemons restart after exit, but a live stalled process need not exit; record doctor and journal output before an explicit systemd restart. Recheck health and send a small request afterward. See daily operation for restart limits, backup, and removal.
Report a useful bug
Include the source revision, OS/architecture, deployment route, redacted command or Nix configuration, expected and observed behavior, doctor output, and a bounded log excerpt. For bridge/effect problems include the request ID and commit SHA. Report at GitHub issues.