Control Tower v0.2.1

Troubleshooting

Start with the flight. Model, HTTP API and A2A responses carry x-ct-flight-id (MCP responses, /v1/models and token counting don't; an MCP refusal carries flight_id in its text). Search for it under Flights to see which key made the call, where it was routed, which gate decided it and what the provider answered. Server logs are on stdout (docker logs controltower); --detailed_debug or CT_LOG_LEVEL=debug adds detail.

Asking for help? Attach a support bundle. It holds the version, settings by name, database state, health and the last day's error counts, and never includes keys, prompts or names: docker exec controltower node dist/server.mjs --support-bundle > support-bundle.json. See SUPPORT.md.

The agent gets an error

Every refusal has a code and a message meant to be read by the agent (or its model). On model calls and HTTP APIs, a key's rate-limit 429 also carries retry-after.

Model calls

/v1/…, /gemini/… and /bedrock/…, in the error envelope of the API the client speaks.

CodeStatusWhat it meansWhat to do
invalid_api_key401No key, or not one Control Tower knowsSend the agent's ct_sk_… key as Authorization: Bearer (OpenAI SDKs: OPENAI_API_KEY; Claude Code: ANTHROPIC_AUTH_TOKEN). A provider's own key (sk-… from OpenAI) won't work unless it was brought over
key_disabled, key_expired401The key was disabled or has expiredKeys → enable it, or create a new one
invalid_request400The body isn't what the API expects: no model, no messages or input, a method the pass-through doesn't coverThe message names the field
model_not_found404No connected provider serves that nameConnect the provider that offers it, check the spelling, or pin it as provider/model. With CT_AUTO_MODELS=0, add the model under Models
model_not_allowed403The key's allowed models don't include itWiden the key's allowed models, or use a model it allows
customer_blocked403The call's customer is blocked under Ledger → CustomersUnblock the customer, or send another one (tags and customers)
region_not_allowed403The call asked for a region (x-ct-region) outside the key's regionsAsk for a region the key allows, or widen them (regions)
region_not_available403No deployment of the model is in a region the key allowsAdd a deployment in the region, or set the region on the model
no_deployment_for_tags403Every deployment of the model is reserved for tags this call doesn't carrySend a matching x-ct-tags, or tag a deployment default
endpoint_not_supported400The model's provider has no such endpoint here (images, audio, moderations or completions on a provider without them)Use a model from an OpenAI or OpenAI-compatible provider
context_window_exceeded400By Control Tower's estimate, the prompt fits no deployment of the modelShorten it, or give the model a context-window fallback
policy_denied403A gate blocks this path. Also: an approver denied the call (Denied by an approver), or an approval ticket was sent with a different request (scope mismatch)The message is the gate's reason (or its name) and the error carries rule_id; find the gate on the Airspace or in Flights. After a scope mismatch, send the original request with the ticket, or ask again
approval_required403Held for a human and nobody answered in time. ct.status is pending with a ticket, or expiredApprove it in the Tower, then retry the same call with x-ct-approval: <ticket> — see Approvals. A retry while the card is still open gets approval_required again, pending
content_blocked400An inspect gate found something it blocks — a secret, personal data, prompt injectionThe message says which detector; remove it from the request, or change the gate's action to mask or flag
rate_limit_exceeded, too_many_parallel_requests429The key's rate or concurrency limit, or a gate's limit (the message names the gate)Slow down, or raise the limit
deployment_busy429Every deployment of the model is at its own rate or concurrency limitRetry shortly, or add deployments
budget_exceeded429A key, team, project or customer budget is used upThe message names the budget; raise it on the Ledger or wait for the period to reset
delegation_required, delegation_loop, delegation_too_deep, delegation_invalid403A call between agents: no valid delegation token, a loop, a chain over 8 agents, or a token that can't be renewedSee Agents calling agents
shutting_down503The instance is restartingRetry shortly (another instance, if you run several)
provider_auth_error502The provider rejected the stored credential (not the agent's key)Providers → Test connection; update the credential
provider_misconfigured502The provider's stored settings are incomplete (Bedrock without access keys, Vertex without a project)Complete them under Providers
provider_model_not_found502The provider says it has no such model, after any fallbacksCheck the deployment's upstream model name
provider_rate_limited, provider_error, provider_overloaded, provider_timeout, provider_unreachable429, 502, 502, 504, 502The provider failed, after any retries and fallbacksAdd a fallback with an alias; set up an outage alert. provider_unreachable is a network error: check HTTPS_PROXY and DNS
provider_stream_error—The provider failed after the stream started; it arrives as an error event in the streamRetry; see the flight for the provider's message
provider_bad_request400The provider rejected the request as invalid; it is not retried elsewhereThe message is the provider's; fix the request (a parameter or input the model doesn't accept)
provider_context_window_exceeded, provider_content_policy400The provider refused the prompt as too long, or refused its contentAdd a prompt too long or content refused fallback
400 on /v1/responses mentioning previous_response_id400The model's provider has no Responses API, so Control Tower translates the call through Chat Completions and can't continue a stored responseSend the whole conversation in input (store: false), or use an OpenAI model. OpenAI built-in tools (web search, file search, computer use) are dropped for these providers

MCP tool calls

A refused tool call is a tool result with isError: true, so the model can read it. Its text ends with JSON: ct_status, flight_id, and details. The flight records the code in the second column.

ct_statusFlight codeWhat it meansWhat to do
deniedpolicy_denied, tool_not_allowed (reason: key_not_allowed), delegation_…A gate, the key's allowed_mcp, an approver's deny, a scope mismatch, or a delegation checkThe text gives the reason. Widen allowed_mcp, or change the gate
not_foundtool_not_foundNo registered server has that toolCheck the name (server__tool on /mcp) and MCP servers
rate_limitedrate_limit_exceededThe key's or a gate's rate limit; retry_after_ms says how longSlow down, or raise the limit
pending, expiredapproval_requiredHeld for a human; pending carries a ticketOnce approved, repeat the call with _meta.ct_approval (or x-ct-approval) set to the ticket — see MCP
content_blockedcontent_blockedAn inspect gate blocked the arguments or the resultChange what is sent, or the gate's action
upstream_errorupstream_…The MCP server itself failedCheck the server under MCP servers → Test

Some problems are JSON-RPC errors instead: -32001 (HTTP 401) for a missing or unknown key, -32004 (404) for an unknown server slug, -32601 for a method that isn't there, -32602 for a tool name without server__, -32000 while restarting, -32603 (internal_error in Flights) for anything unexpected. GET /mcp answers 405: Control Tower doesn't open server-initiated streams, so use POST.

HTTP APIs

/http/<slug>/… refusals are {"error": {code, message, ct_status, flight_id}} with the code in an x-ct-status header too.

CodeStatusWhat it meansWhat to do
missing_api_key, invalid_api_key, key_expired401No Control Tower key, or an unknown, disabled or expired oneSend it as x-ct-key (or Authorization: Bearer ct_sk_…)
api_not_found404No enabled HTTP API has that slugCheck HTTP APIs
path_not_allowed400The path leaves the registered base URL (dot segments, encoded dots)Use a path under the base URL
tool_not_allowed403The key's allowed_mcp doesn't cover <slug>__<route>Widen allowed_mcp
policy_denied403A gate, an approver's deny, or a scope mismatchAs for model calls
approval_required403Held and not answered in time; the ticket is in the body and the x-ct-approval-ticket headerOnce approved, repeat the request with x-ct-approval: <ticket>
approval_pending403A retry with a ticket nobody has decided yet; carries retry-afterWait, then retry with the same ticket
content_blocked400, or 403 for a responseAn inspect gate blocked the request or the responseChange what is sent, or the gate's action
rate_limit_exceeded429The key's or a gate's rate limitSlow down, or raise the limit
upstream_timeout, upstream_unreachable504, 502The API didn't answer in time, or couldn't be reachedCheck the API's base URL and timeout
response_too_large502The API answered with more than 10 MBAsk for less (paging, filters)
shutting_down503The instance is restartingRetry shortly

A2A agents

A2A refusals are JSON-RPC errors; the reason in data is the code in capitals: AGENT_NOT_FOUND, TOOL_NOT_ALLOWED, POLICY_DENIED, APPROVAL_REQUIRED (with a ticket), APPROVAL_PENDING, CONTENT_BLOCKED, RATE_LIMIT_EXCEEDED, SHUTTING_DOWN, the DELEGATION_… codes, INVALID_WEBHOOK, INVALID_AGENT_RESPONSE, UNEXPECTED_STREAM, UPSTREAM_TIMEOUT, UPSTREAM_UNREACHABLE. A stream that fails partway is recorded as stream_error, and a task the agent reports failed or rejected as agent_task_failed / agent_task_rejected. See A2A troubleshooting.

The console and the admin API

CodeStatusWhat it meansWhat to do
setup_code403First-run setup without the right setup codeCopy it from the server's log (docker logs <container>), or open the Open link printed there — see First-run setup
already_setup409Setup has already been doneSign in
rate_limited429Too many sign-in or setup attempts (CT_LOGIN_RPM a minute per email, twice that per address)Wait a minute (retry-after says how long)
unauthenticated401No session, or it ended (7 days, or 12 hours unused)Sign in again, or send the admin key as a bearer token
csrf403A change without the x-ct-csrf headerSend the token from GET /admin/api/me
password_change_required403Signed in with a one-time passwordChoose your own password first
forbidden403Your role is approver or viewerAsk an admin to change it under People
internal_error500Something unexpected; the message has the request idFind the id in the server's log, and include it in an issue

Setup and deployment

Streaming stops, or approvals time out, behind a proxy. Reverse proxies buffer responses and cut long requests. Turn buffering off for Control Tower and allow at least the approval hold time (20 s by default) plus model latency — for nginx:

location / {
  proxy_pass http://controltower:4000;
  proxy_buffering off;
  proxy_read_timeout 600s;
  proxy_http_version 1.1;
  proxy_set_header Upgrade $http_upgrade;     # the console's live updates use a WebSocket
  proxy_set_header Connection "upgrade";
}

Everything is gone after an upgrade. /data wasn't on a named volume, so the new container started with an empty one. Always run with -v controltower-data:/data (or a platform disk mounted at /data). See Install.

"CT_MASTER_KEY differs from /data/master.key. Refusing to start". Both are set and they disagree; stored credentials are encrypted with one of them. Keep the one your data was written with and remove the other.

A console password is lost. An admin resets it under People (a new one-time password). If the only admin lost theirs, reset it with the admin key: see People and roles.

The setup page asks for a setup code. It is printed in the server's log at start, under Setup code, with an Open link that fills it in (docker logs <container>; on a platform, its deploy logs). It is there so that whoever reaches a new install first can't create its admin. See First-run setup.

The master key is lost. Stored provider, tool-server and channel credentials can't be decrypted. Start with a new key and enter those credentials again — keys, gates, zones and history are not encrypted and are kept.

The container can't write /data. The image fixes the ownership of a root-owned volume at startup. If your platform runs containers with a fixed non-root user and read-only ownership, make the volume writable by uid 1000 (node).

Providers time out behind a corporate firewall. Set HTTPS_PROXY (and NO_PROXY for internal hosts) so Control Tower's own calls go through your proxy — see Install.

Port 4000 is taken. Publish another host port (-p 8080:4000) or set --port / CT_PORT / PORT.

--config or --policy stops startup. The log says what's wrong: a file that can't be read or parsed, or a policy naming an agent, model or zone that doesn't exist. Nothing is half-applied; fix the file and restart.

Console and map

The console can't sign in with admin and the master key. The account is created from the admin key only when the key is set — CT_ADMIN_KEY in the environment, or master_key in the config file — and the username is UI_USERNAME (default admin).

Links in Slack, email or webhooks point at localhost. Set CT_PUBLIC_URL to the address people use to reach the console (it's detected on Render, Fly.io and Railway).

The map is empty. Nothing has called the gateway yet — check the key's Connect panel, which turns green on the first request — or a filter is on (All / Active / Gateway / Outside). Fit brings every station into view.

A line is dashed. That traffic was reported by the agent's SDK or OpenTelemetry, not routed through the gateway, so gates can't stop it. Bring it inside on the station shows how to route it. See What is enforced.

The demo won't start. It refuses when your setup already uses one of its model or tool names, so demo traffic can never reach a real provider. The message lists the names.

Alerts

A channel shows failing. The reason is under the channel on the Alerts page; Send test tries again.

  • Slack / webhook: the URL was revoked or is unreachable from the server.
  • Email: Invalid login — wrong username or password (Gmail and Microsoft 365 need an app password); a timeout — the port is blocked or wrong (587 with STARTTLS, 465 with TLS from the start); 554 No SMTP server — set one on the channel or CT_SMTP_URL.

An alert didn't fire. Check the rule's trigger and condition (N times within M minutes), whether it's paused, and its cooldown: after firing, a rule waits and then sends one digest.

Still stuck

Search the issues or open one with the version (controltower --version), how you run it, and the flight id or log lines. Report security issues privately — see SECURITY.md.