Observed on cyrus-ai 0.2.71 with @anthropic-ai/claude-agent-sdk 0.3.260, macOS (Apple Silicon), local subprocess transport.
Symptom: /status reported busy for 3 d 4 h after the last session had completed, and one claude SDK child survived at 0% CPU for the same period. pm2 restart cyrus cleared both.
Cause, from the packaged source:
EdgeWorker.computeStatus() (EdgeWorker.js:1786-1803) returns busy while any runner's isRunning() is true.
ClaudeRunner.isRunning (ClaudeRunner.js:882) is set once in startWithPrompt (:333-337) and cleared at three sites: normal completion (:672), the catch (:683), and stop() (:876). All three exist.
- They are unreachable when
for await (const message of this.activeQuery) (ClaudeRunner.js:597) parks: the child stops producing messages without exiting. There is no timeout on that path (no setTimeout in ClaudeRunner; the SDK's maxIdleMs applies only to the remote sessions transport).
stop() clears the flag synchronously but the child is reaped only by ProcessTransport.close() (reached by unwinding the query) or process.on("exit"). A parked await reaches neither, so an abort() leaves a live child.
AgentSessionManager.getAllAgentRunners() (AgentSessionManager.js:1057) applies no age filter, so a parked runner is polled by computeStatus() indefinitely.
Timeline of the instance: session resumed 2026-09-10T20:08:07Z; no further events; session_stop_requested 2026-09-14T00:53:42Z with no session_stopped after it (that event is only emitted from the catch), i.e. the abort did not unwind the await; process tree showed the child until the restart.
Asks:
- An idle timeout on the local transport's message stream (or in the runner's loop) that aborts a session whose child has produced nothing for N minutes, so the clear sites are reachable.
abort() reaching ProcessTransport.close() even when the stream has not settled, so the child is SIGTERM'd on stop.
Happy to test a patch against this host; the supervisor here now flags both conditions independently (busy with no session record written in 15 min; a child older than the newest record).
Observed on cyrus-ai 0.2.71 with @anthropic-ai/claude-agent-sdk 0.3.260, macOS (Apple Silicon), local subprocess transport.
Symptom:
/statusreportedbusyfor 3 d 4 h after the last session had completed, and oneclaudeSDK child survived at 0% CPU for the same period.pm2 restart cyruscleared both.Cause, from the packaged source:
EdgeWorker.computeStatus()(EdgeWorker.js:1786-1803) returnsbusywhile any runner'sisRunning()is true.ClaudeRunner.isRunning(ClaudeRunner.js:882) is set once instartWithPrompt(:333-337) and cleared at three sites: normal completion (:672), the catch (:683), andstop()(:876). All three exist.for await (const message of this.activeQuery)(ClaudeRunner.js:597) parks: the child stops producing messages without exiting. There is no timeout on that path (nosetTimeoutin ClaudeRunner; the SDK'smaxIdleMsapplies only to the remote sessions transport).stop()clears the flag synchronously but the child is reaped only byProcessTransport.close()(reached by unwinding the query) orprocess.on("exit"). A parked await reaches neither, so anabort()leaves a live child.AgentSessionManager.getAllAgentRunners()(AgentSessionManager.js:1057) applies no age filter, so a parked runner is polled bycomputeStatus()indefinitely.Timeline of the instance: session resumed 2026-09-10T20:08:07Z; no further events;
session_stop_requested2026-09-14T00:53:42Z with nosession_stoppedafter it (that event is only emitted from the catch), i.e. the abort did not unwind the await; process tree showed the child until the restart.Asks:
abort()reachingProcessTransport.close()even when the stream has not settled, so the child is SIGTERM'd on stop.Happy to test a patch against this host; the supervisor here now flags both conditions independently (busy with no session record written in 15 min; a child older than the newest record).