Skip to content

feat(ai,agent,coding): visible model retries and interrupted-stream recovery - #3

Merged
rsbin1178 merged 4 commits into
mainfrom
feat/stream-recovery
Oct 7, 2026
Merged

rsbin1178 merged 4 commits into
mainfrom
feat/stream-recovery

Conversation

@rsbin1178

Copy link
Copy Markdown
Owner

What this changes

A model call that dies mid-stream used to kill the whole interaction, and the retries that did
happen were invisible. This branch makes both visible and survivable, following the strategies
Claude Code and Grok Build actually ship.

1. Model retries are announced (ai, internal/coding, TUI)

  • ai.StreamRetry carries a RetryNotice{Attempt, MaxRetries, Delay, Reason}; the retry middleware
    emits one per replay, and Collect ignores it, so a notice is never counted as output.
  • New Coding event model.retry, kept as live reducer state (State.Retry) that the next event of
    any kind clears; it is not durable state.
  • The activity row renders Retrying… · retry 2/10 · in 4s · connection error, with a countdown
    driven by the existing activity clock.

2. A turn whose stream broke after output is re-issued (agent, internal/coding)

WithStreamRecovery(attempts, base, maxDelay) re-issues such a turn after emitting
CandidateDiscarded, so a frontend drops the partial output it already rendered. Only the loop can
retract what a consumer has seen, which is why this is not the middleware's job; a failure that
produced nothing stays with the middleware, so the two budgets never stack on one failure. Nothing
from a failed attempt is committed and no tool ran, so a replay cannot duplicate content or repeat
an effect. Coding runs share one policy: ten re-issues, doubling from 2s, capped at 30s.

3. Budgets and counters aligned with the shipped frontends (ai, coding/model, TUI)

  • Ten retries for a request that produced nothing (was five), keeping the 500ms base, full jitter,
    the 30s cap, and a provider Retry-After.
  • Certificate validation failures are no longer retryable; transient TLS conditions still are.
  • The notice counts retries, not tries (attempt / max_retries), mirroring the retry_state
    payloads Grok Build pushes to its UI.
  • Waiting for the model… · no data for 25s replaces the thinking row once an open turn has carried
    no model progress for twenty seconds, so a gateway that buffers no longer looks like a slow model.

Research behind the numbers

Recorded in .trellis/tasks/10-07-stream-retry-recovery/research/provider-retry-strategies.md:

  • Claude Code: "retries transient failures up to 10 times with exponential backoff",
    CLAUDE_CODE_MAX_RETRIES default 10 (capped at 15), a 20s no-data banner, byte/stream watchdogs,
    a first-byte deadline retry once, and certificate failures reported on the first attempt.
  • Grok Build: a live retry_state update; a real transcript from this machine shows
    {"attempt":1,"max_retries":15,"reason":"reqwest error stream: Transport error: error decoding response body","error_type":"http"} — the same truncated-SSE class this branch recovers from —
    plus models.max_retries / rate_limit_retry_threshold / subagent_rate_limit_max_attempts
    knobs and Retrying (attempt n/m) in the UI.

Verification

  • go test ./... green (89 packages) with TERM=xterm-256color; go vet ./... clean; gofmt clean.
  • New tests: notice counters and backoff schedule, recovery on/off/budget-exhausted/pre-output and
    non-retryable paths, an end-to-end runtime test (provider truncates after a delta → the run
    recovers, reports one notice, discards the candidate, commits the re-issued answer), reducer
    lifecycle including the batched delta path, projector mapping, and the TUI retry/stall rows.
  • golangci-lint could not run here: the installed binary was built with Go 1.26.4 and panics on
    the local Go 1.27.1 standard library, including on untouched packages.

Out of scope

A transcript marker for discarded partial text, journaling retry notices, config knobs for the
budgets, and an aborting idle watchdog (the stall row reports silence without aborting).

…stalled

A stream that dies before its first event is replayed by the model
middleware, but nothing reported it: the wait happened silently, and the
only way to notice was that the turn looked slow.

Announce every re-attempt on the model stream (ai.StreamRetry carrying the
attempt about to start, its backoff and a short cause), project it to the
model.retry Coding event, keep it as live reducer state, and render it as
the activity line "Retrying… · attempt x/y · in Ns · reason".

The notice is progress rather than output: Collect ignores it, it never
counts as produced content, it cannot close an open text or reasoning part,
and it is not durable state — the next event of any kind ends the wait.
A mid-stream truncation (io.ErrUnexpectedEOF from a dropped SSE body)
failed the whole interaction even though the turn contributed nothing: no
message was committed and no tool ran, so replaying it could neither
duplicate content nor repeat an effect.

WithStreamRecovery re-issues such a turn, bounded, after emitting
CandidateDiscarded so a frontend drops the partial output it already
rendered. That retraction is why the loop owns this case and the model
middleware does not; a failure that produced nothing stays with the
middleware, so the two budgets never stack on one failure.

Coding runs (interaction, subagent, team worker, draft and proposal
agents) share one policy: ten re-issues, doubling from two seconds up to
thirty, which covers a few minutes of outage before reporting it.
… frontends

The retry behavior now follows what Claude Code and Grok Build do, with the
numbers taken from their documentation, their binary, and live retry payloads:

- Coding runs get ten retries for a request that produced nothing (was five),
  keeping the 500ms base, full jitter, the 30s cap, and a provider Retry-After.
- Certificate validation failures are no longer retryable: the caller has to
  fix them, so they surface on the first attempt, while transient TLS
  conditions such as a handshake timeout stay retryable.
- The notice counts retries rather than tries: attempt is the 1-based retry
  ordinal and max_retries is the budget, mirroring the fields Grok Build pushes
  to its frontends; the TUI reads "retry 2/10".
- The activity row reports a silent model — "Waiting for the model… · no data
  for 25s" — once an open turn has carried no model progress for twenty
  seconds, so a gateway that buffers a response no longer looks like a slow one.
- Give the parent-stream progress switch a default branch, which is how this
  repository keeps event-type switches forward compatible.
- Assert the rejected retry payloads with require, so the table stops at the
  first failure instead of reporting once per row.
- Drop the unused streamRecovery.enabled helper; the loop already spells the
  same condition out as a remaining-budget check.
@rsbin1178
rsbin1178 merged commit c2aa9f0 into main Oct 7, 2026
4 checks passed
@rsbin1178
rsbin1178 deleted the feat/stream-recovery branch October 8, 2026 01:43
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant