Skip to content
github-actions[bot] edited this page Oct 8, 2026 · 33 revisions

clauth daemon: headless scheduler + status feed

clauth daemon runs the same background refresher the TUI runs (spawn_refresher) with no UI: refresh usage, rotate tokens ahead of expiry, run the fallback chain's auto-switches, and publish ~/.clauth/status.json for external readers. Codex profiles ride the same loop: their usage poll, their chain refresh and their own chain's switches (Codex). clauth status --json prints the same shape single-shot, no daemon required. This document is the read contract for both.

Scope note: by default the daemon carries no external surface at all: it publishes a file and reads a file, and everything reaches it through ~/.clauth. --listen opts into one, a TLS REST API that only paired devices can use and whose mutations are a switch and the fallback-chain edits (see below); without that flag the daemon itself listens on nothing (an adopted shunt gateway it runs is its own process on its own port, below). What holds either way: anything that needs eyes (a diverged live login, a manual Claude Code /login, an unprovable identity) is refused and logged, and the TUI stays the only resolution surface. The REST switch refuses those cases with a 409 exactly as the scheduler and the MCP tool refuse them.

Process model

  • Singleton: one advisory lock (~/.clauth/clauthd.lock) held for the process lifetime, with the holder's pid written to an unlocked sidecar (~/.clauth/clauthd.pid) for ps-level diagnosis. Informational: the flock, never the number, proves presence.
    • The pid stays out of the lock file because Windows locks are mandatory: a --status reader in another process could not read bytes inside the daemon's exclusive lock.
    • A second clauth daemon exits 0 by default (already running (pid <n>)), so a spawner that fires repeatedly can't pile up idle daemons (#57).
    • --standby opts into take-over: it parks and takes over the moment the holder exits, for a supervisor's instance queueing behind a manually run one under launchd KeepAlive{SuccessfulExit=false} (which never restarts a clean exit).
    • The standby queue is one deep: a second flock (clauthd-standby.lock) holds the slot; any further instance exits. --no-standby is the default's explicit spelling, kept for callers already passing it.
    • A dead holder's flock auto-releases, so a supervisor with restart-on-crash keeps exactly one scheduler alive without pidfile bookkeeping.
    • The TUI header's always-present [ daemon ] chip reads this lock (presence) plus status.json freshness (green = fresh feed, amber = stalling, dim = no daemon) to show whether one is running.
  • Asking before spawning: clauth daemon --status prints running (pid <n>, feed fresh|stale[, standby waiting]) and exits 0 when a daemon is up; with none up, exit 1 and nothing on stdout.
    • It creates nothing, so a menu-bar app or a wrapper script can gate its spawn on the exit code instead of starting a process to find out.
    • A lock file it cannot test at all (no working flock, e.g. some NFS/CIFS mounts) is its own failure with the io error attached, never an exit 1 that reads as "none running" and sends a supervisor into a respawn loop.
    • The header chip answers the same question by dimming instead, which is why the two read the lock through different paths. The default already exits the moment it loses the race, which suits the same callers.
    • --standby is the one to keep out of a pure supervisor unit: the supervisor is the sole starter and wins the race alone, so a standby only earns its keep when a manual run and a unit coexist.
  • Replacing for an upgrade: clauth daemon --replace terminates the running daemon and takes over. It reads the holder's pid sidecar and confirms the pid is still a running clauth daemon by its argv, so another clauth subcommand sharing the binary name is never signalled.
    • It SIGTERMs the holder and waits for the flock to auto-release on death; on a timeout it escalates to SIGKILL, then claims (on Windows there is no graceful kill for a console process, so both passes are taskkill /F). A pid it can't confirm bails rather than signal blind.
  • Starting and stopping from the TUI: the action menu's start daemon runs plain clauth daemon (no --listen) from ~/.clauth, its stdout and stderr appended to ~/.clauth/daemon.log, detached so it outlives the TUI and its terminal: on unix it leads its own process group, on Windows it runs on a hidden console of its own and leaves the TUI's job object where that job allows it. stop daemon sends --replace's termination and claims nothing after it; once no daemon is left it also stops a gateway and any proxy the stopped daemon left running (every Windows stop, where taskkill /F gives the daemon no chance to), recording each stop in that child's marker like a daemon start would. When another daemon holds the lock right after the stop (a parked standby, a supervisor's restart), that one's start reclaims them instead.
  • Probes take the lock they read, briefly: both the header chip's daemon_health and --status try-lock a free file and release it. A starting daemon therefore re-tries a lost race (3 attempts, 100 ms apart) before it accepts that another instance is up.
    • A real holder keeps its lock for life, so anything that clears on a retry was a reader.
  • Watchdog: a wedged tick can freeze the single-threaded loop. The cross-process state flock a tick may block on is capped at 25 s, so a flock-blocked tick times out and retries rather than hanging.
    • If no tick completes in 30 s at all, the daemon abort()s for a clean supervisor restart, freeing the usage lease.
    • A suspended process (Modern Standby, S3, a VM pause) freezes the loop but not the clock; the watchdog charges a poll that overshot its window no running time, so the heartbeat gap is judged on time the loop actually had — a suspend/resume never reads as a stalled tick, and a loop that wedges before or after the resume still aborts.
    • A legit keychain switch sits inside both margins: on macOS it reads the Keychain item and writes it back, each of those killed at 10 s.
    • Everything one flock hold spends in security is capped at 20 s no matter how many reads and writes it makes.
  • Log hygiene: every daemon-visible stderr line carries an ISO-8601 UTC prefix, enabled only in daemon mode. An interactive terminal instead diverts its lines to ~/.clauth/clauth.log so a background thread never paints over the TUI; a redirected or piped stderr keeps the bare line.
    • With --listen on, every request adds one clauth api: line — peer address, the device that asked (- for none), a bounded method+path summary, the response status — so a client's poll rate sets the log volume: a tray polling twice a minute writes two lines a minute for the daemon's life. The size cap below trims that like anything else.
    • ~/.clauth/daemon.log is size-capped in place when a supervisor, or the TUI's start daemon, points stderr at it. The in-place trim is only sound for an APPEND-mode fd: use launchd StandardErrorPath or systemd StandardError=append:....
    • A non-append redirect (file:, a plain >) keeps its own offset, so the next write after a trim leaves a sparse NUL hole and the size cap is defeated. The daemon checks its own stderr at boot and warns loudly when it is a non-append file, so a defeated cap shows up in the log instead of only in this page.
  • Usage-history samples: the lease holder appends every live /usage reading to ~/.clauth/profiles/<name>/usage_history.jsonl, the series the burn rate and burn-aware switching read. The daemon keeps it advancing with no TUI open.
    • Retention is 2 days, re-trimmed on a 6-hour cadence rather than only at startup, so a long-lived daemon bounds the file without a restart.
  • The 5h alert: with bell_command set in profiles.toml (Configuration), the daemon runs it once each time an account's 5h figure reaches that account's bell_threshold, and again only after the figure falls back below. It judges the figures it just built for status.json: only a Fresh, non-stale entry moves the alert, so an outage's frozen number neither fires nor re-arms it, and a disabled account that is not the active one never fires. A restart above the line fires once more. A program that fails to start is logged and the crossing still counts as fired. With the TUI open too, its toast and the command both fire.
  • Single usage fetcher (usage-fetch.lock lease): every instance (the daemon and each open TUI) runs the same refresher, but only the one holding the usage-fetch.lock flock fetches usage, rotates tokens, and decides switches.
    • The rest hydrate from the shared disk caches instead of double-polling the usage API, double-rotating the single-use refresh chain, or re-deciding switches.
    • The lease is first-come and held for the process lifetime, no preemption, so the switch-decider never thrashes between processes; a waiter takes it over within one tick of the holder exiting (flock auto-release).
    • The daemon normally boots first and holds it, but a TUI already fetching keeps the lease until it closes, and the daemon then hydrates while still publishing status.json every tick.
  • The shunt gateway: the Services tab's shunt card adopts a config into ~/.clauth/gateway.toml and toggles its disabled flag (Interface and keys); until one is adopted the daemon runs no gateway and the feed's gateway reads absent. Once a config is adopted and disabled is off, the running daemon keeps one shunt run process alive beside it, never inside it: it restarts one that exits (1 s backoff, doubling to 60 s), never starts one while anything else listens on the gateway's port, and publishes what it sees as the feed's gateway object (below).
    • Only the singleton holder runs it; a --standby instance starts it once promoted.
    • The TUI header shows a [ shunt ] chip left of [ daemon ]: green while the gateway serves; amber while it starts, restarts, stops or fails its health check; red while it refuses to run until its setup is fixed (the Services tab's shunt card says why); dim while nothing runs it (none adopted, disabled, held off, or no daemon). When the TUI finds a gateway adopted and enabled with no daemon running, at launch or when the daemon goes away, it warns once: the shunt gateway will not run, or a shunt gateway the last daemon left still runs while a gateway a killed daemon left behind keeps serving; start daemon in the a menu starts one.
    • A config edit to a setting shunt reads only at startup restarts the running gateway — a normal stop, then a fresh start — once shunt check accepts the config as it stands, whichever way the setting changed. Those settings are [server] bind, max_concurrent_requests and shutdown_timeout_seconds; [server.limits] max_request_header_bytes and max_url_length; anything in [server.access_control], [server.rate_limits], [sentry] or [otel]; [server.spend] state_path; [server.pool] usage_refresh_seconds and state_path; [server.status] refresh_seconds; whether [server.admin] (the admin-table fix's confirm says the gateway restarts once to load it.), [server.gateway], [server.spend], [server.codex_endpoint], [server.usage], [server.oauth_usage] or [server.status] exist; and whether [server.status] sources is empty. daemon.log names each changed setting, never its value. A moved line, a new comment or a TOML respelling of the same value restarts nothing; a value TOML keeps distinct but shunt reads as equal (1 and 1.0), or a setting written out at shunt's default where it was absent (an empty sources aside), restarts once. No other config edit restarts it. A config that does not parse as TOML, or that gives one of these tables a non-table value, restarts nothing and logs nothing until it is fixed. A refused check keeps the running gateway and is re-run when the config's bytes change, or 60 s after the last refusal; a refusal that repeats the last one logged stays out of daemon.log.
    • The gateway's stdout and stderr go to ~/.clauth/gateway.log, size-capped like daemon.log, never into daemon.log, which would drown in shunt's per-request lines. An env-file line that assigns nothing is named in daemon.log by its number, never its text, the first time a start skips it.
    • SIGTERM, SIGINT and SIGHUP stop the gateway first (one SIGTERM, up to 4 s), then the daemon dies of the same signal. A signal the daemon inherited as ignored stays ignored, so a nohup daemon still survives a hangup. A gateway still draining after those 4 s keeps running with its stop deadline in ~/.clauth/gateway-child.json, and the next daemon start waits out the rest of it and kills what is left.
    • Under systemd, give the unit KillMode=process: control-group hands the gateway a second SIGTERM, which makes shunt skip its drain, and mixed kills it outright when the daemon exits.
    • Windows has no stop signal for a console process: the gateway outlives the daemon, and the next start stops it, or the TUI's stop daemon once no daemon is left.
  • The clauth proxies: every enabled row of ~/.clauth/proxies.toml runs its own clauth-<service>-proxy process beside the daemon, never inside it, through the same supervisor that runs the gateway (register one with clauth proxy enable <service>, then the running daemon picks it up). The daemon re-reads the registry each round, so a new enable starts without a daemon restart, and a disabled or removed row stops its child through the supervisor's own round.
    • Only the singleton holder runs them; a --standby instance starts them once promoted.
    • Each proxy's stdout and stderr go to ~/.clauth/proxies/<service>/clauth.log, size-capped like daemon.log, never into daemon.log. Its admin token is ~/.clauth/proxies/<service>/clauth-admin-token, handed to the child only as CLAUTH_PROXY_ADMIN_TOKEN_FILE.
    • A foreign listener on a proxy's port starts nothing beside it: the feed reads foreign and names what answered, until the port clears.
    • On a signal the daemon stops the gateway and every proxy with the same discipline: one SIGTERM each, then the one aggregate stop budget. A proxy still draining after that keeps running with its stop deadline in ~/.clauth/proxies/<service>/clauth-child.json, and the next daemon start waits out the rest of it and kills what is left (the systemd KillMode=process note above applies to every proxy the same way).
    • Windows has no stop signal for a console process: a proxy outlives the daemon, and the next start stops it, or the TUI's stop daemon once no daemon is left.

Supervising the daemon on Windows

clauth ships no supervisor. A macOS launchd unit with KeepAlive restarts the daemon, and a Linux systemd unit does the same; Windows has no built-in supervisor for a plain user program, so a daemon that exits stays down and ~/.clauth/status.json freezes at its last write.

Two ways to keep it up:

  • Run it from a spawner that restarts it whenever the process ends (a loop that re-runs clauth daemon each time the child exits; a crash releases the singleton lock, so the next run claims it cleanly).
  • A Scheduled Task that fires clauth daemon every few minutes: the singleton lock makes the redundant start exit 0 with already running (pid <n>), so repeating the task while one is up changes nothing.

clauth daemon --status answers whether one is running without spawning anything (exit 1 = none up), so a spawner can gate on it.

Either way, a reader of the frozen feed detects a dead daemon from generated_at: the stamp is rewritten every tick, so a stamp much older than refresh_interval_ms means the daemon is gone, stuck, or the machine was asleep. Show the last-known numbers with a stale cue, never as current truth.

REST API (--listen)

clauth daemon --listen 0.0.0.0:8443 also serves the feed and the switch over HTTPS, for running the daemon on one machine and a client (clauth-tray) on another. Off unless the flag is passed. TLS comes from this host's lego certificate, on all three platforms clauth supports — macOS, Linux, and Windows. Only two things differ by platform: where that certificate lives, and how the host's own FQDN is discovered. Both are spelled out below, and --cert/--key skip the derivation entirely for hosts where it cannot work.

  • The address is optional. A bare clauth daemon --listen binds 0.0.0.0:8443, the spelling this page uses everywhere else; pass an explicit ADDR:PORT to bind anything narrower (--listen 127.0.0.1:8443 for a loopback-only listener behind a reverse proxy, say). The shorthand deliberately defaults to every interface rather than loopback, because a listener the remote client cannot reach is not a useful default for this flag. Nothing listens unless --listen is passed either way — the default fills in the value, never the flag.

  • TLS is not optional. There is no plaintext mode and no flag to ask for one, because a device's bearer token crosses the connection on every request. The certificate is read once at startup from the platform's lego directory (see the next bullet), in lego's own naming: <fqdn>.crt, <fqdn>.issuer.crt, and <fqdn>.key. Any certificate the issuer file adds that the leaf file already carries is dropped, since lego usually writes the whole chain into <fqdn>.crt and a chain that repeats a certificate is malformed. A renewal reaches the listener on the next daemon restart. A missing or unreadable file, or a port already taken, refuses to start rather than running a daemon that looks healthy while the remote client stays dark.

  • Where the certificate lives, and how the host is named, are the only two things that differ across platforms:

    Default certificate directory Host FQDN from
    macOS, Linux /etc/lego/certificates hostname -f
    Windows %AppData%\lego\certificates (C:\Users\<you>\AppData\Roaming\lego\certificates on a stock install) PowerShell [System.Net.Dns]::GetHostEntry($env:COMPUTERNAME).HostName

    The directory is a default, overridable through tls.json (next bullet). The FQDN is not overridable at all — it is whatever the rest of the box already believes it is called, and a certificate issued to a name the host does not answer to would be the actual bug. Point lego at the Windows directory with its --path flag: lego's own default there is .lego under the working directory, which is useless for a daemon, since the working directory is whatever started it. The Windows default is per-user, matching where clauth keeps everything else (~/.clauth) — so a daemon run as you finds it, while one run as a Windows service under LocalSystem would resolve a different %AppData% and needs tls.json pointed at a directory both accounts can read. Note that Windows' whoami /fqdn is not the lookup and could not be one: its "FQDN" is the current user's Active Directory distinguished name (CN=…,DC=…), not the machine's, and it fails outright for a local account.

  • ~/.clauth/tls.json points at the certificates. Written with this platform's default the first time a --listen daemon starts, so there is a real file to edit rather than a documented path to retype:

    {
      "schema": 1,
      "cert_dir": "/etc/lego/certificates"
    }

    Edit cert_dir and restart to serve from anywhere else — a Windows box whose lego lives outside %AppData%, a Linux one following a distribution's own packaging convention, or a staging certificate in a scratch directory. The file is never rewritten once it exists, so an edit survives every restart and every upgrade. A tls.json that cannot be parsed — or that carries an empty cert_dir — refuses to start instead of falling back to the default: quietly serving out of a directory you believe you moved away from is a worse failure than not starting, and the error names the file. Delete it to get the platform default back. A schema newer than this build knows is read anyway (the field it needs is a path either way), so a downgrade does not strand your configured directory.

  • The certificate has to be this host's own. Whichever lookup applies, the certificate the daemon serves must be issued for that exact FQDN — and clients must reach it by that name, https://<fqdn>:8443. A bare IP fails certificate verification no matter what is listening, which is the easy mistake to make once --listen defaults to 0.0.0.0: binding every interface says nothing about which names the certificate covers. If the box's own name disagrees with the name you got the certificate for, the daemon does not fall back or guess — it looks for files that are not there and refuses to start. Two consequences worth planning for: the host needs a real, publicly resolvable FQDN for Let's Encrypt to validate at all, and a client on a LAN needs that same name to resolve to the daemon's reachable address (split-horizon DNS or a hosts entry), since the name in the certificate, not the address you dialed, is what gets checked.

  • --cert and --key skip the derivation entirely. For the hosts where it cannot work rather than merely points somewhere else: on a tailnet node hostname -f answers a name no certificate covers, and tailscale cert <machine>.<tailnet>.ts.net writes a .crt and a .key with no issuer file beside them and none of lego's naming. A --listen address in Tailscale's ranges (100.64.0.0/10, fd7a:115c:a1e0::/48) with neither flag and no lego certificate for this host refuses to start naming that exact tailscale cert command and the two flags to pass, ahead of lego's missing-file error. 100.64.0.0/10 is also carrier-grade NAT space, so the message names the range rather than calling the host a tailnet node. Passing both by path reads exactly those two files — no hostname -f, no tls.json, and no sibling .issuer.crt even if one happens to sit there, so the chain served is the one in the file you named. The pair is required together and only alongside --listen; a lone --cert is a usage error rather than a silent fall back to lego. Everything else is unchanged: clients still dial the name the certificate covers, and the certificate is still read once at startup, so the restart hook below applies to a tailscale cert renewal too.

  • Renewal is yours to automate. Let's Encrypt certificates are valid for 90 days, and clauth does not renew them — it only reads what lego left on disk, once, at startup. So a --listen daemon needs two automated steps, not one: a periodic lego ... renew (cron or a systemd timer on macOS and Linux, a Scheduled Task on Windows, or whatever else supervises the host), and a daemon restart afterwards, since the certificate is never re-read while the process lives. Renewing without the restart is the failure mode to watch for — the daemon keeps serving the now-expired certificate it read at boot, and every client rejects the handshake while the files on disk look perfectly current. clauth daemon --replace --listen is the in-place restart for that hook — --replace composes with --listen, but the hook has to pass both, since a --replace on its own takes over without a listener. One thing that hook cannot recover from: if some OTHER process (not the daemon) holds the port, --replace terminates the daemon first and then dies at the bind, leaving the host with no daemon at all — refresh and auto-switch stop until it is started again. The renewal hook is the one place this bites by construction: whatever binds the port while the daemon is being restarted (a second daemonless listener, a port-hijacking probe) is a squatter the incumbent daemon's death does not clear. A start under --replace --listen is fatal in one more set, after the incumbent is already gone: the squatter bind failure just named, an unreadable certificate under --cert/--key (read above the claim, so it fails before anyone is terminated), and a legacy-token import (below) that cannot write the device list, which leaves auth_token.json in place for the next start to retry. Any other failure shape is fatal without a --replace too; this is the extra set --replace adds. Check clauth daemon --status after the hook, or make the hook fall back to a plain restart on failure. Nothing warns you in advance: a certificate that expires under a running daemon produces client-side TLS errors, not a clauth log line.

  • Devices. Every request but a pairing authenticates as one named device, and each device holds a tier fixed on this machine when its code or token was minted: view reads the feed, control may also switch accounts. A device joins in one of two ways: clauth devices pair <name> [--control] [--sessions] prints a one-time code for the device to redeem over the API, and clauth devices add <name> [--control] [--sessions] mints a token here and prints it once, for a client you configure by hand; --sessions (which needs --control) is the trusted-machine grant that lets the device create sessions through the API once [serve] session_creation is on. A token is 32 CSPRNG bytes through SHA-256, hex, so 64 characters. clauth devices lists the devices (--json for an array; each row carries sessions, true when the device holds the grant), clauth devices allow-sessions <name> grants the sessions flag to a control device later (a device without the control tier is refused: revoke it and re-pair with --control), and clauth devices revoke <name> removes one, grant included; its next request gets a 401 with no restart, because ~/.clauth/devices.json is read per request. That file (0600) holds each device's name, tier, sessions grant, join time and a SHA-256 of its token, never the token, so a copy of it lets nobody in; the comparison is constant-time over digests, so neither a token's length nor its first wrong byte is observable in the timing. A device with a tier this build does not know, one a newer clauth paired, authenticates and is refused every route that needs a device with 403 device_tier_unknown, and the other devices are unaffected.

  • Pairing. clauth devices pair <name> prints an 8-character code as XXXX-XXXX, drawn from the OS random source over Crockford's base32 alphabet (0-9 and A-Z without I, L, O, U), then waits. The code is valid for 5 minutes, redeems once, and is dropped after 5 wrong tries; a new pair replaces a code still waiting, and the replaced wait exits 1 saying so. The device posts it to POST /api/v1/pair (below) and gets its token in the answer. Only the code's SHA-256 reaches the disk, in ~/.clauth/pairing.json (0600). The code prints alone on stdout and the wait reports on stderr: paired '<name>' (<tier>) and exit 0, or exit 1 with the reason (dropped, expired, replaced, or a stdout reader that left before the code printed, which withdraws the code). Ctrl-C withdraws a code still waiting and exits 130; a code that already paired, or was dropped, expired or replaced, reports that ending and its exit code instead. With --control the code is a control credential while it is live: whoever enters it first gets control. Only this host's daemon serving --listen can redeem it, and pair warns when no daemon is running at all.

  • Upgrading from the single token. Before pairing, the API had one bearer token in ~/.clauth/auth_token.json. The first --listen start after the upgrade that binds its port imports it as the control device legacy and deletes the file, so the client holding it (clauth-tray) keeps working on the same bytes. A file carrying any tier but control is neither imported nor served, and a file holding no usable token is not imported; daemon.log says which, once per start. If a downgraded clauth later mints a fresh auth_token.json, the next upgraded start hands legacy that token instead, since it is the one the client was given last. A token revoked after its import is never imported again, even from a file the import could not delete: the device list keeps the SHA-256 of the last token it imported. --print-token and --rotate-token are gone: a digest cannot be printed back, and a rotation is clauth devices revoke legacy followed by a new pair or add.

  • Every route but the pairing needs a device's token, health included. The slot is taken at accept(), before the TLS handshake and before any token is seen, because refusing at that point is the cheapest refusal available and the alternative is spawning threads for unauthenticated peers without bound. So an unauthenticated client occupies a slot while it is connected, and a 401 or a failed pairing closes the connection rather than keeping it alive, so it cannot hold one, and every wrong pairing code costs its sender a new connection. What bounds the occupancy is the clock, not the token: a client that connects and says nothing gets only the 10s first-request timeout, not the 120s lifetime, and one that says something unauthenticated is answered and closed at once.

Route Needs What it does
GET /api/v1/health view {"ok":true,"version":"<ver>","schema":2}. schema is the status feed's, so a reader can refuse a daemon newer than it knows.
GET /api/v1/status view The status.json body below, byte for byte off disk. Conditional: the ETag digests everything in the body except generated_at, so a feed rewritten by a tick that changed nothing answers 304 with no body. ?wait=<secs> on a request already carrying the current tag holds it open until the accounts actually move (capped at 60s, inside the connection's own 120s lifetime), which is how a client follows a switch without polling. ?all=1 rebuilds the body to include disabled accounts, which the published file always hides, never waits, and is still conditional off its own built body's tag. With no file on disk yet (the daemon's first tick has not landed) or one caught mid-replacement, the body is built here from the same caches instead — the single-shot shape — and carries its own ETag the same way.
GET /api/v1/events view One server-sent-events stream, the rest of the connection. It opens with retry: 1000, then event: herdr — data: {"present":true,"panes":[…]} with one {"pane_id","workspace_id","agent","agent_status"} object per herdr pane (the snapshot every connect, reconnect and re-subscribe hands over, so no agent-status change is missed across a gap) or {"present":false} when no herdr socket answers — then event: status frames. Each status frame is the whole published feed, the status.json bytes off disk exactly as GET /api/v1/status serves them, one data: line per line of the file (the client re-joins them with a newline), id: set to the same ETag, one frame each time the tag moves (a tick that moved only generated_at is not a change); before the first tick lands it sends nothing until the file appears. event: pane_agent_status frames relay herdr's pane.agent_status_changed events byte for byte; a pane opened or closed re-subscribes on a fresh herdr connection and hands over a new snapshot. A : keepalive comment is written after 15 s of silence. The stream ends cleanly at the connection's 120 s lifetime — no Content-Length, Connection: close — so a reconnecting client gets the current feed and snapshot first again. herdr's socket is resolved once per stream by one herdr status server --json probe, and the daemon targets herdr's default session wherever it was started — every herdr call it makes strips the HERDR_* session variables and passes no session, so a daemon started inside a herdr pane or under systemd serves the same panes either way, and an inherited HERDR_SOCKET_PATH is never consulted; a socket that is absent, refuses, or drops flips to {"present":false} once per transition and the stream keeps serving status frames, retrying the socket every 5 s while it lives; a stream that could not resolve a socket at all stays absent until its reconnect. A browser's native EventSource cannot send Authorization: Bearer, so a web client reads the stream with a fetch-based SSE reader that honours retry:.
HEAD /api/v1/health, HEAD /api/v1/status, HEAD /api/v1/events, HEAD /api/v1/openapi.json, HEAD /api/v1/panes, HEAD /api/v1/sessions, HEAD /api/v1/sessions/<id>, HEAD /api/v1/gateway, HEAD /api/v1/proxies view The same status line and headers their GET would answer, with no body and Content-Length: 0 — a deliberate deviation from RFC 9110 §8.6, which would state the length the GET sends. A HEAD /api/v1/status?wait= carrying the current If-None-Match long-polls exactly like the GET; the wait is not skipped for a HEAD.
GET /api/v1/openapi.json view The OpenAPI document for this API. clauth daemon --dump-openapi writes the exact same bytes to stdout and exits 0 without starting a daemon, so CI pins the spec without a listener.
POST /api/v1/switch control Body {"profile":"<name>"} (resolved case-insensitively, against the Claude Code roster alone: a codex name answers 404 profile_not_found, and the chain routes below edit the Claude Code chain alone). Returns {"ok":true,"previous":…,"active":…}. Republishes status.json before answering, so every reader parked on GET /api/v1/status?wait= is woken by the same switch rather than by the next tick.
POST /api/v1/chain/order control Body {"members":["<name>",…]}, the full new order (each name resolved case-insensitively; a member the roster no longer holds drops out). Returns {"ok":true,"members":[…the order that landed, canonical names…]}. Republishes status.json before answering, like the switch.
POST /api/v1/chain/threshold control Body {"profile":"<name>","threshold":<number>}. Sets that chain member's fallback_threshold (finite, 0..=100). Returns {"ok":true,"profile":"<canonical name>","threshold":<number as stored>}. Republishes before answering.
POST /api/v1/chain/wrap-off control Body {"wrap_off":true|false}. Sets the chain-global quota spent behaviour (wrap_off in the feed). Returns {"ok":true,"wrap_off":<bool>}. Republishes before answering; setting the value it already holds still answers 200 and republishes. The feed's fallback.armed (the member the switch has active) is the switch route's to set, never a chain route's.
POST /api/v1/pair no token Body {"code":"<code>"}; case, dashes and whitespace do not matter, and O, I and L read as 0, 1 and 1. Returns 201 {"ok":true,"name":…,"tier":…,"token":…}, the one response that ever carries a token.
GET /api/v1/panes view Every herdr pane on this host joined to the clauth sessions running inside it, on process ids only — the tokens.clauth tag is display and never joins. A row belongs to a pane when its pid is the pane's foreground process group or a process listed in its process-info; its kind is session when that pid leads the group (the pane's own clauth start or clauth resume) or is a listed clauth running start or resume under a wrapper, and delegate for any other listed process (a claude child, a clauth mcp in flight). foreground_process_group_id is null when the pane's process-info did not answer (it closed between the two calls) or when the pane has no foreground job, and the pane's cwd is null when herdr has none. On Windows herdr lists only the pane's agent root, so no session joins there. Herdr absent (not installed, or no server answered on its socket) answers 200 with herdr.present: false, the one fixed sentence naming the state, and panes: [] — never an error. At most 1 + N herdr calls at 2 s each, where N is the pane count. Each pane also carries agent_session_id: the agent's own session id herdr detected in the pane (its agent_session of kind id), the id GET /api/v1/sessions/<id> pages, null when herdr detected none.
GET /api/v1/sessions view A page of the Claude Code transcript index across the shared store and every live isolated store, newest first (updated desc, then id), previews redacted exactly as clauth sessions --json redacts them. ?limit= 1..=200 (default 50); ?before= the previous page's next_before, an opaque cursor (null on the last page that carries rows). Each row: id (the transcript stem), last_ran_profile (null when unknown), workspace (empty when the transcript records no cwd), updated (ISO-8601 +00:00), first_message/last_message (null when none), store (global or isolated:<profile>). No tokens/cost: the full-parse figures are not served. Only the rows on the page are opened (a bounded head and tail read each); the walk itself opens nothing. A malformed before or limit is 400 bad_request.
GET /api/v1/sessions/<id> view A page of one transcript's records, verbatim: {"ok":true,"id":…,"records":[{"offset":<byte offset>,"record":<the JSONL line as a JSON object>},…],"next_before":<offset>|null,"malformed":<count>}, oldest first within the page, the last limit records (1..=500, default 100) whose lines end at or before ?before= (a byte offset; absent or past the end = the file's end; a before inside a line consumes only the lines that end before it). next_before is the start offset of the oldest line the page consumed, a record it served or a line it counted as malformed, and null when that line starts at byte 0 or the page consumed nothing. A line that is not a JSON object (a torn last line mid-write, a blank line, a JSON array) is skipped and counted in malformed, never served and never an error. A page holds at most 4 MiB of records, except that a single record larger than that is served alone, never cut. clauth parses no record: each is Claude Code's own line, handed over as written, so a Claude Code record change never needs a clauth release; clients model the shapes they render. The id is percent-decoded once and looked up among the transcripts the listing shows, never joined onto a path: an unknown or path-shaped id, and a transcript that cannot be opened, answer 404 session_not_found.
POST /api/v1/panes/<id>/prompt control Body {"text":"<non-empty, no NUL>"}. Runs herdr agent prompt <id> <text> without --wait on herdr's default session (the HERDR_* session variables stripped, like every daemon-side herdr call); 200 {"ok":true} says herdr accepted the submission, and the events stream carries what the agent does with it. The pane id in the path is percent-decoded once (a generated client sends w1N%3Ap19); an id not shaped like herdr's <workspace>:<pane> (each half an alphanumeric, then alphanumerics, - or _) answers 404 pane_not_found before herdr is asked. 400 bad_request (no non-empty text, a NUL in it, or not an object), 404 pane_not_found (herdr's agent_not_found or pane_not_found), 409 agent_blocked (the agent is waiting on a prompt of its own; answer it with keys or the terminal stream), 503 herdr_unavailable (herdr not installed, or its server not running: the /panes sentences), 502 herdr_refused for any other herdr refusal, whose output goes to daemon.log and never the body. The audit line records the pane and the text's length, never the text.
POST /api/v1/panes/<id>/keys control Body {"keys":["y","enter"]}, 1 to 32 herdr key names in press order, each 1 to 32 characters of letters, digits, -, + and _ (esc is the canonical Escape; herdr decides which names exist). Runs herdr pane send-keys <id> <key>… on the default session; 200 {"ok":true}. Answers as the prompt route minus the 409: a name herdr does not know is a 502 herdr_refused. The audit line records the pane and the key count, never the names.
POST /api/v1/sessions control + the device's sessions grant + [serve] session_creation Creates a herdr tab in a directory and starts an agent in its pane, so the app can open a session it then drives through the stream, prompt and keys routes. Body {"cwd":"<absolute, existing directory>","profile":"<clauth profile>"?,"kind":"<herdr agent kind>"?,"workspace":"<herdr workspace id>"?}: profile and kind are exclusive; neither means bare claude; a profile (codex profiles included) runs clauth start <profile> in the pane; a kind runs herdr agent start … --kind <kind>. Three gates, in order, before anything is created: [serve] session_creation in profiles.toml is read fresh per request and refused 403 session_creation_off naming the key while off; the calling device must hold the sessions grant (403 sessions_grant_required naming clauth devices allow-sessions <name>); the device must hold the control tier (the table's own 403). Then herdr tab create --cwd <cwd> --no-focus [--workspace <id>] on herdr's default session, then the agent; the answer carries workspace_id, tab_id, pane_id and agent_status (null for a profile launch, whose readiness the events stream carries later). 400 bad_request names the field (a relative or missing cwd, both profile and kind, a malformed kind or workspace; a body that is not JSON answers a bare 400 bad_request); 404 profile_not_found (the name is in neither the claude nor the codex roster); 409 workspace_not_found (the named workspace, or the default one when none is named, does not exist); 503 herdr_unavailable (herdr not installed, its server not running, or a tab create that did not answer inside 2 s — then a tab may exist unattributed, and daemon.log says so); any other herdr refusal after the tab exists closes it first (best-effort, a failed close logged with the tab id) and answers 502 herdr_refused — or 404 pane_not_found for a refused pane run — with herdr's output in daemon.log, never the body; a hung agent start is killed at 65 s and closed the same way (a profile launch is bounded at 2 s, it returns as soon as the command is handed to the pane). One audit line per creation names the device, the agent form, the cwd, the tab and the pane. No creation cap: a control device with the grant can already start any number of agents by typing into a shell pane, so a cap here would bound nothing.
GET /api/v1/panes/<id>/stream view to watch, control to type A WebSocket (RFC 6455, version 13) bridging herdr's terminal stream for that pane; the pane id in the path is percent-decoded once, like every <id> route. The device tier decides the half you get, never a request parameter: view is bridged onto herdr terminal session observe and control onto terminal session control. Server→client frames are herdr's own newline-delimited JSON records verbatim (terminal.frame renders, terminal.closed ends the stream — a second controller refused, a displaced incumbent, all of it); client→server frames are herdr's control commands, `{"type":"terminal.input","text"
GET /api/v1/gateway view The managed shunt gateway: the gateway object of the feed below, read from the daemon's supervisor per request rather than off status.json, so it can be one tick ahead of the file. Read-only: no route changes the gateway.
GET /api/v1/proxies view The managed clauth proxies: the proxies array of the feed below, read from the daemon's supervisors per request rather than off status.json, so it can be one tick ahead of the file. Read-only: no route changes a proxy.

POST /api/v1/switch failures: 404 profile_not_found, unknown profile · 409 switch_refused, because the target is disabled, its credentials were rejected by a refresh (including a transient refresh failure, which the reason names and which may succeed again on retry), the live login is one clauth has not saved, or the profile disappeared between the config read and the switch itself (the body's reason names the fix, and nothing was changed) · 409 switch_in_progress when another switch is still running · 503 state_locked when another clauth process is holding the state lock, so the same request will work shortly · 500 switch_failed, an unexpected failure whose reason is a fixed sentence pointing at daemon.log for the full error chain · 400 bad_request, a malformed body. Refusals carry {"ok":false,"error":"<code>","reason":…}.

POST /api/v1/chain/order failures: 400 chain_order_invalid when the list is not a permutation of the current chain (a missing, extra, duplicated or unknown member; the reason names which) · 400 bad_request, a malformed body · 409 edit_in_progress when another chain edit or a switch is still running · 503 state_locked, as for the switch · 500 edit_failed, a disk error whose reason is a fixed sentence pointing at daemon.log.

POST /api/v1/chain/threshold failures: 404 profile_not_found, unknown profile · 409 not_a_member when the profile is stored but not in the chain (the reason says to add it on the Fallback tab first) · 400 bad_request, a malformed body or a threshold that is not finite or outside 0..=100 · 409 edit_in_progress, 503 state_locked, 500 edit_failed, as for the order route.

POST /api/v1/chain/wrap-off failures: 400 bad_request, a malformed body · 409 edit_in_progress, 503 state_locked, 500 edit_failed, as for the order route.

POST /api/v1/pair failures: 403 pairing_refused for every failed redemption (a wrong code, one dropped after its 5 wrong tries, an expired one, or none waiting) with one fixed reason telling the user to run clauth devices pair again, so the answer says nothing about which it was · 400 bad_request, a body that is not {"code":"…"} or a code that is not 8 alphabet characters once normalized, which costs no attempt · 503 state_locked and 500 internal as for the switch. Any answer but a 201 closes the connection.

Every route but the pairing can also answer before its own work: 401 unauthorized (with WWW-Authenticate: Bearer) when no paired device holds the bearer, a path in no row of the table included; 403 control_required when a view-only device calls a route that needs control; 403 device_tier_unknown when the device carries a tier this build does not know; and 500 internal when devices.json cannot be read at all (announced once in daemon.log, not per request).

The switch is the same action clauth <name> and the MCP tool perform, so it inherits their gates rather than reimplementing them, and the daemon's main loop picks the result up on its next tick through the ordinary reload path.

Connections persist. HTTP/1.1 keeps the connection open unless the client sends Connection: close (HTTP/1.0 needs Connection: keep-alive to opt in), so a polling client handshakes once rather than once per poll. Every response states its own Connection: disposition and carries a Content-Length, and a kept-alive one advertises Keep-Alive: timeout=<seconds left>, max=<requests left>: the figures are what remains of this connection's budget, not a nominal maximum, so a client is never told it has time it does not have. Pipelined requests are served strictly in order, one at a time.

What makes that safe is that framing is never ambiguous: Content-Length is the only accepted framing, chunked transfer encoding is refused outright, two Content-Length headers that disagree are a hard error, and any framing error answers once and closes rather than trying to resynchronize. Those are the conditions request smuggling needs, and none of them is available here.

Limits: 8 KiB of headers, 64 KiB of body, 100 requests and 120 seconds per connection, a 10s deadline on any single read or write, and 32 concurrent connections. The 10s and the 120s are separate bounds and mean different things. A read that times out while the connection is idle between requests is just an idle connection, so the wait resumes up to the 120s budget; the same timeout part-way through a request is a peer trickling bytes to hold a slot, and fails with 408. A connection that has not sent its first request gets only the 10s, since a client that connects and says nothing has not earned the idle allowance. CLAUTH_NO_API=1 disables the listener whatever the flags say, for killing it without editing the unit that passes --listen.

Neither bound is a response-time budget, and a switch is the case where that matters. The events stream is the one answer that consumes the whole lifetime on purpose: it runs until the connection's 120 s and then closes cleanly, so a client sizes its reconnect against the same budget; the chain routes wait on the same flock as the switch and answer 503 state_locked the same way. The 10s applies to one read or one write, not to the handler between them: POST /api/v1/switch waits up to 25 seconds for the cross-process state flock (clauth's own bound, sized around a macOS Keychain switch) and may run a token refresh after that, so a switch can legitimately take tens of seconds before a single byte of the response is written. Nothing times out — the whole of it sits inside the connection's 120s lifetime — but a client that sets a 10s socket deadline because this page mentions 10s will give up on a switch that was about to succeed. Size a client timeout against the 120s, and treat a 503 state_locked (state lock held) as the retryable answer rather than an abandoned request.

Readers should follow the same evolution rule as the feed (below): ignore unknown fields, and refuse only on a schema greater than what they know.

~/.clauth/status.json

Written each daemon tick, and by whichever process lands a switch when no daemon is running to do it — so the published active_profile is never an account the operator has already switched away from, whether the switch came from the TUI, clauth <name>, or the MCP tool. A running daemon owns the file: a switch made elsewhere leaves the write to that daemon's next tick (≤1s), which republishes it with the scheduler's live fetch_status / next_refresh_at / pending_switch that a single-shot build cannot see. A switch the daemon itself performs through POST /api/v1/switch, and a chain edit through POST /api/v1/chain/*, are the exception: they republish before answering, so a client waiting on the feed is woken by the switch rather than by the tick after it. A switch republished with no daemon running keeps the stamp the daemon last wrote, or the epoch when no daemon has ever published or the feed on disk cannot be read or parsed: a fresh stamp is how a reader tells a live daemon from a dead one, and the switch-side write publishes the new active_profile without forging that liveness — it carries the stamp forward, never mints one. Atomic (tmp + rename into place), 0600. Never carries a token, secret, or key: names, tiers, percentages, timestamps only.

A reader that wants a switch the moment it happens has three ways in: GET /api/v1/events streams every feed change as it lands. Over the network, GET /api/v1/status?wait=<secs> with the current ETag blocks until the feed's content actually changes — no polling interval to lose time to. On the same machine, watch this file and ~/.clauth/profiles.toml, which every switch surface rewrites and nothing else touches that often.

Compare more than the modification time when you do. A daemon rewrites this file every second through tmp + rename, two writes can land inside one filesystem clock granule, and a feed republished for a different active account is often exactly as long as the one it replaced — so (mtime, len) can miss a real switch. The renamed inode always differs, which is what makes the change detectable whatever the clock did.

{
  "schema": 2,
  "generated_at": "2026-09-07T19:04:40+00:00",
  "active_profile": "kitty",
  "pending_switch": null,
  "wrap_off": false,
  "active_codex_profile": "cx",
  "codex_fallback_chain": ["cx"],
  "codex_wrap_off": false,
  "refresh_interval_ms": 300000,
  "clauth_version": "0.15.2",
  "gateway": {
    "state": "healthy",
    "config": "/home/kitty/.config/shunt/shunt.toml",
    "binary": "shunt",
    "port": 3001,
    "pid": 48213,
    "version": "0.49.1",
    "answerer": null,
    "floor": "0.48.0",
    "restarts": 0,
    "last_exit": null,
    "reason": null,
    "since": "2026-09-07T18:02:11+00:00"
  },
  "proxies": [],
  "profiles": [
    {
      "name": "kitty",
      "active": true,
      "rolling_token": false,
      "provider": "anthropic",
      "base_url": null,
      "tier": "Max 5x",
      "harness": "claude",
      "has_live_session": true,
      "auth_status": "ok",
      "fetch_status": "Fresh",
      "stale": false,
      "fetched_at": "2026-09-07T19:04:20+00:00",
      "next_refresh_at": "2026-09-07T19:09:20+00:00",
      "auto_start": true,
      "auto_start_queue": { "position": 1, "next_open_at": "2026-09-07T21:34:20+00:00" },
      "bell_threshold": 90,
      "fallback": { "position": 1, "threshold": 95.0, "armed": true },
      "windows": [
        { "label": "5h",      "utilization_pct": 42.0, "resets_at": "2026-09-07T23:00:00+00:00" },
        { "label": "7d",      "utilization_pct": 18.0, "resets_at": "2026-09-12T17:00:00+00:00" },
        { "label": "7d Opus", "utilization_pct": 30.0, "resets_at": "2026-09-12T17:00:00+00:00" }
      ],
      "third_party": null
    },
    {
      "name": "cx",
      "active": true,
      "rolling_token": false,
      "provider": "openai",
      "base_url": null,
      "tier": "plus",
      "harness": "codex",
      "has_live_session": false,
      "auth_status": "ok",
      "fetch_status": "Fresh",
      "stale": false,
      "fetched_at": "2026-09-07T19:04:31+00:00",
      "next_refresh_at": "2026-09-07T19:09:31+00:00",
      "auto_start": false,
      "auto_start_queue": null,
      "bell_threshold": null,
      "fallback": null,
      "windows": [
        { "label": "5h", "utilization_pct": 12.0, "resets_at": "2026-09-07T22:10:00+00:00" },
        { "label": "7d", "utilization_pct": 40.0, "resets_at": "2026-09-13T09:00:00+00:00" }
      ],
      "third_party": null
    }
  ]
}

Field notes

Field Semantics
schema Integer, currently 2. Bumped ONLY on a breaking change; additive fields do not bump it (evolution rule below).
generated_at Write stamp, ISO-8601 UTC with an explicit +00:00 offset (all timestamps are; parse the offset, never key on a Z suffix; the writer does not emit one). Readers derive staleness from it: a stamp much older than refresh_interval_ms means the daemon is gone/stuck, so show last-known data with a stale cue, never spin. A feed republished by a switch outside the daemon carries the daemon's last stamp forward, or the epoch when no daemon ever published or the feed on disk cannot be read or parsed, so staleness keeps meaning daemon-absent rather than following whichever process last wrote the file.
active_profile The profile whose credentials are currently installed, else null.
pending_switch A switch the daemon has accepted but not yet applied ("<name>"), else null. This remains the target name alone even when the internal queue also records why a rejected api key caused the move; that cause exists only to drop the move if the key is repaired before dispatch. Exists so readers can show in-flight truth instead of a timing heuristic. Always null from the single-shot CLI.
wrap_off The fallback chain's stop-vs-stay-on-active flag, verbatim from state.
active_codex_profile Additive (schema stays 2; absent from an older writer ⇒ null). The codex profile the codex chain anchors on, else null. active_profile, wrap_off and pending_switch above stay the Claude Code slots and never name a codex profile; a codex switch is never pending, it lands in state at once and takes effect at the next clauth start (Codex).
codex_fallback_chain Additive (absent ⇒ []). The codex chain in walk order, verbatim from codex-profiles.toml.
codex_wrap_off Additive (absent ⇒ false). The codex chain's wrap_off.
clauth_version Additive (absent ⇒ ""). The version of the clauth that wrote the feed. An empty codex roster from a daemon that predates codex and one from a daemon with no codex profiles are otherwise byte-identical; this is how a reader tells them apart.
gateway Additive (schema stays 2; absent from an older writer ⇒ null). The managed shunt gateway, the object GET /api/v1/gateway serves: paths, a port, a pid, versions and states, never an env value, the admin token or the shunt config's content. With no daemon (clauth status --json, or a feed a switch republished) it is read off ~/.clauth/gateway.toml alone: the record's own verdict, or unobserved for a gateway that would run.
gateway.state One of absent (no gateway.toml), disabled, held (the TUI's stop shunt: off until its start shunt or the next daemon start; a disabled record still reads disabled), no_config (the adopted config is gone), yaml_refused (the record names a YAML config), misconfigured (reason says why: the record, its env file or its bind does not read, the config cannot be inspected, gateway.log does not open, or the binary does not run), binary_missing (binary names what was looked up), foreign (something clauth did not start listens on the port; nothing starts beside it, and the port is probed again every 5 s), starting, healthy, unhealthy (running, but /health stopped answering, or never answered within 10 s of the start; never killed for it), below_floor (it reported a version below floor, so clauth stopped it and waits for gateway.toml or the binary to change), restarting (it exited on its own; the next start waits out the backoff), stopping, unobserved (no supervisor looked). Changing the adopted config, the binary or the env file in gateway.toml stops the running gateway and starts it afresh at once; a config edit to a setting shunt reads only at startup restarts it once shunt check accepts it (the gateway bullet above); an edit inside the env file is read at the next start only.
gateway.config, gateway.binary The adopted config, and the shunt it runs as (the record's path, else shunt looked up on PATH). Both null when no record was read (absent, or misconfigured from a record that does not load) and while a gateway a previous daemon left running is being stopped; binary is also null on yaml_refused.
gateway.port, gateway.pid The port the gateway binds, or the one a foreign listener holds; the pid of the gateway the daemon started while it runs, or of the one it is stopping. null where none is known.
gateway.version, gateway.answerer version is what /health last reported: the gateway's own, or a foreign listener's when it answers the way shunt does. It is whatever text that listener sent, so render it as untrusted. answerer is set on foreign alone: shunt, not_shunt (an HTTP answer that is not shunt's) or no_answer (it took the connection and never answered).
gateway.floor The oldest shunt clauth supervises, "0.48.0".
gateway.restarts, gateway.last_exit restarts counts the restarts after an exit clauth did not ask for, since the daemon started; a stop clauth asked for (a gateway.toml change, disabled, a version below the floor) is not one. last_exit is how the last gateway process ended, { "code": integer | null, "signal": integer | null } (signal always null on Windows), null before any has.
gateway.reason Why a misconfigured gateway cannot start, else null.
gateway.since ISO-8601 UTC of the last change of state; null until a supervisor has set one.
proxies[] Additive (schema stays 2; absent from an older writer ⇒ []). The managed clauth proxies, one object per ~/.clauth/proxies.toml row in service order, the array GET /api/v1/proxies serves: a service, a port, a pid, versions and states, never a token, an env value or a file's content. With no daemon (clauth status --json, or a feed a switch republished) it is read off the registry alone: each row's own verdict, or unobserved for a proxy the daemon would run. A registry that does not read while a daemon runs keeps publishing the live slots it holds instead of [], and logs the failure once per distinct error by each of two readers of the registry.
proxies[].service The proxy's service, the <service> in its binary name, the [<service>] table key.
proxies[].state One of absent (no registry row), disabled, binary_missing (binary names what was looked up), no_token (the admin token file is missing; reason names the fix), manifest_refused (reason says why; held until the binary or the row changes), misconfigured (reason says why: the binary or the log cannot be read into a spawn), foreign (something clauth did not start answers on the port; nothing starts beside it, and answerer names what), starting, healthy, unhealthy (running, but /health does not answer, or answers a status other than ok; never killed for it), contract_mismatch (it served a contract major clauth does not speak, a contract that is not MAJOR.MINOR, a /health naming another service, or one naming no service; clauth stopped it and holds off), restarting, stopping, unobserved (no supervisor looked).
proxies[].binary, proxies[].port, proxies[].pid The binary the proxy runs as (the row's recorded path while it exists, else the clauth-<service>-proxy on PATH); the port it binds, or the one a foreign answerer holds; the pid clauth spawned, while one runs. null where none is known.
proxies[].version, proxies[].contract, proxies[].answerer version and contract are what /health last reported: the proxy's own, or a foreign answerer's; render both as untrusted. answerer is set on foreign alone: {"proxy":{"service":…}} (another proxy answered; render its service as untrusted), not_proxy (an HTTP answer that is not a proxy /health) or no_answer (it took the connection and never answered).
proxies[].restarts, proxies[].last_exit restarts counts the restarts after an exit clauth did not ask for, since the daemon started; a stop clauth asked for (a row change, disabled) is not one. last_exit is how the last proxy process ended, { "code": integer | null, "signal": integer | null } (signal always null on Windows), null before any has.
proxies[].reason Why the proxy cannot start or was stopped (no_token, manifest_refused, misconfigured, contract_mismatch), else null.
proxies[].since ISO-8601 UTC of the last change of state; null until a supervisor has set one.
profiles[].provider One of four cases: "anthropic" for a profile with no endpoint of its own (the managed base_url is unset), the recognised provider's display name ("DeepSeek", "Z.ai", …) for a typed third-party account, and "generic" for every other endpoint: an api-key account no provider recognises (litellm, LM Studio, ollama, a LAN gateway). Keyed on the managed base_url alone, so the label never contradicts the base_url published beside it; an operator-authored ANTHROPIC_BASE_URL reroutes requests without changing it. A codex entry publishes "openai", with base_url null.
profiles[].harness Additive (absent ⇒ "claude"). "claude" or "codex": which roster the entry comes from. Codex entries are appended after every Claude Code one, so a reader that predates the field still reads the prefix it always read. Names are unique across both rosters, so name alone still identifies an entry.
profiles[].tier Plan label ("Max 5x", "Pro", "Free"…), null when nothing on disk claims a tier. Opaque display string, so never switch on it. A canceled subscription reports its real post-cancellation tier ("Free") and never the word canceled: cancellation is a status, not a tier, and no field here exposes it. null is rarer than it looks: a token-only account resolves its tier from the OAuth subscription_type alone, Free included. A codex entry carries the ChatGPT plan word the account reports, lowercased ("plus", "pro"), from its last usage poll or, before any, from the login itself; never a Claude label.
profiles[].auth_status "ok" | "expired" | "broken" | "unknown". broken = last refresh rejected as revoked/invalid → excluded from fallback walks, refused as a switch target. expired = access token past its expiry, refresh not yet run (pending or blocked). broken outranks expired. Reports on the credential a profile STORES, not on where its requests route: a hybrid (an OAuth pair kept alongside a base_url) reports expired on a dead token like any other account. Absent ⇒ "ok". A codex entry (harness: "codex") publishes "broken" while its chain is quarantined (Codex), "ok" when it has a usage cache, and "unknown" when it has none; expired has no codex reading.
profiles[].rolling_token bool (additive, schema stays 2; absent ⇒ false). Always false on a codex entry: a rolling token is a Claude Code mechanism. true when the sidecar currently HOLDS a rolling bearer (content-classified — a refresh token present means mis-filled before anything else is considered, then a plan stamp or any scope beyond the setup pair means rolling; the same truth the TUI renders), not when the config flag is merely on: a dead chain degrades the sidecar onto its static mint with the flag still set, and this key follows the sidecar. While true the bearer is plan-stamped, refresh-less, and its expiry is HOURS out. Readers MUST key their token-row rendering off this: a rolling token drawn through the static mint's 30-day warning ramp shows a healthy credential as dying, and a mint drawn through the rolling countdown promises re-stamps nobody will make. (The 2026-07 incident this contract descends from ran the other way entirely — every surface looked healthy while the split's protection was silently off — which is why the key follows what the sidecar HOLDS rather than any flag.) A mis-filled sidecar (rotating pair) publishes false: it is not a rolling token, and the TUI renders it as the DANGER state it is. false likewise for a static claude setup-token mint, whose year-scale countdown is the honest one.
profiles[].fetch_status "Fresh" | "Cached" | "Failed" | "RateLimited" | "AuthExpired": the usage fetch's last outcome, so readers can distinguish live bars from last-known. "AuthExpired" is TERMINAL and means action required rather than a retry pending: this account's usage credential is dead (a revoked api key or an Alibaba Model Studio console session). Only an operator re-entering the key or re-logging in clears it. clauth has stopped polling it, so never read the next_refresh_at beside it as a scheduled attempt. Resolution order, highest first: a live daemon's OAuth store, then its third-party store, then a durable per-profile record of a dead credential, then a derivation from the profile's own cache (Fresh when its last fetch is under one interval old, else Cached), dated off the OAuth body's fetched_at stamp (a plan-only cache rewrite moves the mtime without producing a new reading, so the mtime never re-ages a dated body) and off the cache mtime for an undatable OAuth body, a third-party cache, or a hybrid's OAuth cache. Failed and RateLimited now reach third-party profiles as well as OAuth ones. An api-key profile with a warm cache is never reported as unfetched. A profile nothing has ever fetched stays null (no cache at all), "AuthExpired" included. Never appears on an OAuth profile, whose analogous state is auth_status: "broken". A codex entry derives it from its usage cache alone, live daemon or not: "Fresh" under one interval old, "Cached" past it, null with no cache; "Failed", "RateLimited" and "AuthExpired" never appear on one, and its fetched_at / next_refresh_at come from that cache's write time.
profiles[].stale bool (additive, schema stays 2; absent ⇒ false). true when the reading is distrusted, by either arm: a deep-slot stuck RateLimited — fetch_status == "RateLimited" AND the consecutive-429 streak past the active-retry cap, so the usage throttle never drained and no Fresh read is coming (the same judgment the daemon's auto-switch and Usage row use) — or cache age past 2 × max(refresh_interval_ms, 5min) + refresh_interval_ms, the longest a degraded fetch cadence can leave before the reading is one nobody is maintaining. The stuck arm selects the same source as fetch_status: the live OAuth status + streak when present, otherwise the third-party status + provider-429 streak. A hybrid's OAuth reading therefore wins; a pure provider account can now become stuck/stale after its own deep 429 streak. The chain bypasses the fresh-reading gate for that state, but still requires last-known exhaustion when a window exists; a windowless stuck account is reading-dead because the evidence can never arrive, so it is walked off too. The age arm applies to the single-shot clauth status --json; the stuck arm needs live streak stores and is always false single-shot. A live window pinned at the 100% cap (windows_maxed) that the refresh_spent_accounts opt-out skips is exempt from the age arm: its figure cannot change by polling, so age distrusts nothing about it. A third arm covers a reading nothing can date: an OAuth cache whose body carries no fetched_at, or one stamped in the FUTURE, publishes stale: true with fetched_at null: its windows[] stay, and no field claims to date them. Every OAuth account upgraded from a build before fetched_at existed reads that way until its first live fetch stamps a body. The age arms qualify a FIGURE, so an OAuth profile publishing an empty windows[] (none cached, or every window lapsed) is never stale, whatever its age: there is no number for a reader to discount. That gate is OAuth-only. A third-party (api-key) account's figures — provider-derived windows[] where its provider publishes them, its balance and bars otherwise — have no window exemption, so it does still go stale on cache age. Readers should dim the meter / show a "stuck" cue rather than render the frozen number as current truth. false for a shallow/transient RateLimited, for any cache younger than the threshold, and for an OAuth row with no window to qualify; stale ignores fetch_status, so a Cached or AuthExpired row over an old cache publishes stale: true (the two signals say different things: fetch-status is the last outcome, stale is the age). A codex entry always publishes false: neither arm is judged for it, so date its figures off fetched_at yourself.
profiles[].next_refresh_at ISO-8601 UTC of the next scheduled usage refresh, or null when none is pending. null covers a never-cached profile, a single-shot reading whose derived stamp (the same clock fetch_status derives from, plus the interval) is already past, and, with refresh_spent_accounts off, a spent (100%-capped) account the scheduler skips until its window resets. Treat null as "no refresh scheduled", never as overdue. A non-null stamp is not a promise either: an "AuthExpired" profile is suppressed from the cadence and still carries an ordinary future stamp, so gate on fetch_status before believing one.
profiles[].fallback null when not in the chain; else 1-based position, threshold (%), armed (this member is the active one the auto-switch watches). Always null on a codex entry: the codex chain is published once, top-level, as codex_fallback_chain, and a member's position is its index there.
profiles[].auto_start_queue Additive object { "position": integer, "next_open_at": ISO-8601 UTC | null } (schema stays 2). null when the global queue toggle is off or this profile holds no slot — including while a switch-grade kick block stands against it, since a profile that cannot open a window is not a queue member and does not count toward the shared 5h / N. next_open_at: null means no opening has been observed yet, so the queue is due now; a non-null stamp in the past means the same, since the stamp names the earliest moment the queue may open its next window and one that has passed has cleared the gap. The anchor is every observed opening, not only the daemon's own kicks: a window clauth opened is recorded in usage_history.jsonl as the kick lands, and one opened out of band (a live Claude Code session on that account) becomes the anchor once its samples hold the same boundary long enough to be sure. Single-shot status --json re-derives that from disk on every invocation; the daemon publishes the scheduler's in-memory value, which takes the out-of-band opening up on the next tick that could actually elect someone. Either way next_open_at is an estimate: taking an opening up moves the anchor FORWARD, so for fixed membership the stamp moves later than a reader was last told, and a reader that cached it will be early.
profiles[].fetched_at ISO-8601 UTC, nullable. When the figures were READ, not when the file was written. For an OAuth account it is the body's own fetch stamp, so a cache rewrite that produced no new reading (a plan-only refresh) does not move it; for a third-party account it is the provider cache's write time, whose only writer is a fetch. null means clauth cannot date these figures: no cache, or an OAuth body with no stamp or one stamped in the future. A null here alongside a non-empty windows[] is the undated arm of stale above, and the two travel together.
profiles[].windows[] label is derived, not an enum: "5h" and "7d" always; the third is a plan-tier label ("7d Opus"…). Treat labels as opaque display strings, never keys to switch on. utilization_pct 0-100 float; resets_at nullable. A window whose resets_at has passed is not listed at all: past its reset the figure would be the previous window's last utilization, not a current reading, so the row drops rather than publishing a stale number — a 5h window absent from the array is one that has lapsed, never existed, or lives outside the array (a third-party account whose provider publishes no recognised window), and a row without resets_at is one the endpoint did not stamp, which stays. On a third-party (api-key) profile the array carries that account's provider-derived 5h/7d windows, the two labels only. Never a 30d or a best-effort scan's figure. resets_at is the provider's own reset instant, null when it reports none.
profiles[].third_party { "available": bool } for api-key profiles once probed, else null, including an api-key profile whose provider has never been reached (no cache yet). Plain reachability; structured balances deliberately deferred. Provider WINDOWS publish under windows[], never inside this object.
profiles[] membership A user-disabled account (clauth disable) is excluded from profiles[] by default; no field marks a profile disabled, absence from the array IS the signal. The active profile is always present regardless of its own disabled flag, so active_profile never names a profile missing from profiles[]. Every codex profile is always present (a codex profile cannot be disabled), after the Claude Code entries; on a codex entry auto_start is false and auto_start_queue, bell_threshold and third_party are null, and windows[] carries the 5h and 7d rows its usage poll filled.

Evolution rule (the load-bearing part)

  • Writers: additive only under the same schema: new fields may appear, existing fields never change type/meaning. A breaking change bumps schema. The codex fields (active_codex_profile, codex_fallback_chain, codex_wrap_off, clauth_version, profiles[].harness, the "unknown" auth_status value and the codex entries themselves) are additive under schema 2: a reader that knows none of them sees the Claude Code feed alone, with the codex entries trailing the array. The gateway object and every gateway.state value are additive under schema 2 the same way, and the proxies array with every proxies[].state value is additive under schema 2 the same way again.
  • Readers: ignore unknown fields; default absent optional fields (absent auth_status ⇒ "ok", absent harness ⇒ "claude"); refuse only on schema greater than what they know, showing "daemon newer than me" rather than garbage.

clauth status --json

Same schema, produced single-shot from the on-disk caches with no daemon and no network fetch: pending_switch is always null, generated_at is the print stamp, and freshness/next-refresh derive from each profile's own cache (the OAuth body's fetched_at stamp when it carries one, the cache mtime otherwise). One code path builds both (daemon::status_json::build_status), so the key SHAPE cannot drift between the daemon feed and the CLI snapshot. Values can: the single-shot form derives fetch_status from the caches, so a profile the live daemon shows as Failed or RateLimited reads as Cached here at the same instant. Poll the feed, not the CLI, when the fetch outcome matters. The codex entries and the top-level codex slots are built from the same on-disk files the same way on both paths; nothing live feeds them.

"AuthExpired" is the one exception, exact on both paths. Single-shot it comes from a durable per-profile record written when a fetch died on a credential that cannot self-heal. That record applies only as long as the credential recorded with it still matches the one on disk, so a re-login retires it with no daemon and no timer involved. A profile the daemon has never fetched carries no record and reads null.

clauth status --json --all (or its --disabled spelling, equivalent) is the single-shot way to reveal disabled accounts in profiles[]; the running daemon's own published ~/.clauth/status.json file always hides them (every daemon-side build_status call passes include_disabled: false), so a reader on another machine gets the disabled roster from GET /api/v1/status?all=1 instead of shelling out to the single-shot form, and a reader on this machine has both.

Clone this wiki locally