perf(server): stop loading the whole database on hot orchestration paths - #297
Merged
Merged
Conversation
Session starts, checkpoint capture, and per-event projection loaded full projection snapshots or thread histories. They now use targeted by-id queries, cache materialized settings so secrets are not re-read per token, skip projectors that ignore an event, and cap shell delegation text. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
|
The latest updates on your projects. Learn more about Vercel for GitHub.
|
The shell snapshot now selects open delegations plus the newest terminal ones per parent thread in SQL instead of hydrating every record, and streamed delegation upserts carry the same capped result text. The provider command reactor reads its drain baseline before subscribing so buffered events are not marked as seen. The checkpoint test no longer adds manual Effect runners, and the bot docs describe the card limit. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
|
This comment has been minimized.
This comment has been minimized.
The provider command reactor now retries its subscription until no commit lands between the two sequence reads, so drain neither skips buffered events nor waits for an event the subscription missed. Concurrent settings reads after a change share one secret materialization, and the shared shell reducer keeps the same number of finished delegations per thread as the server snapshot. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
…order The startup retry could drop provider intents that committed between its sequence reads. Keep the first subscription and replay the gap from the event store, skipping live events the replay already covered. The client delegation cap now breaks updatedAt ties like the shell snapshot query. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
… failure The startup gap replay stopped at the event store's default 1,000-event limit, and a failed replay marked the whole gap as handled, so buffered provider intents were filtered out. Replay the full bounded range, retry from the last delivered event, and let the live stream deliver anything the replay did not. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
A replay that failed after retries left drain unable to tell which gap events the live buffer still held. Read the bounded gap eagerly during start, retry transient failures, and fail startup if the store cannot return it, so no intent is dropped and drain never waits on a lost event. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Several server hot paths did work proportional to the entire environment instead of the thread at hand. Starting a Codex or Kimi session loaded every project, bot, thread, message, and activity through a full projection snapshot just to find one thread. Every streamed assistant token re-read secret files from disk to check one settings flag. Each persisted event wrote all 14 projector cursors and reloaded the thread row, and checkpoint capture loaded a thread's full message history about four times per turn. Since node:sqlite is synchronous, these all block the event loop and grow with total chat history.
Session starts now use targeted by-id lookups with new narrow queries for delegations, groups, and bots. Materialized settings are cached and invalidated on writes, so secrets are read once instead of per token. Projectors declare the event types they handle and unrelated ones are skipped, with all cursors advancing in one batched write. Checkpoint capture uses the lightweight checkpoint-context query, ingestion caches the per-thread runtime context across content deltas, and the shell snapshot keeps open plus the 20 most recent finished delegations per thread with result text capped at 2,000 characters. Two behavior boundaries: the session-start lookup excludes archived and deleted threads, which the old snapshot included, and a thread whose project row is missing no longer gets a checkpoint.
Verified with 401 passing tests across the touched orchestration and provider suites, including new tests for the settings cache, projector skipping, and checkpoint linkage, plus 241 server, decider, and routine tests and a clean workspace typecheck. Fixing a pre-existing race in the ProviderCommandReactor test drain was required: the faster pipeline exposed that drain could return before subscribed events reached the worker.
Created with Claude Fable 5.1 in Claude Code.
🤖 Generated with Claude Code