You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
When the model bound to an agent profile stops on a provider rate limit or quota (for example a Codex/OpenCode usage-limit error), Kandev treats it as a session failure that needs a human, not as the model going away. The profile's configured fallback model is never used.
Current behavior, verified in this repository:
Profile fallback is start-model-only.StartModelPolicy (FallbackModel, AutoFallback, RequireExactModel) in apps/backend/internal/agent/runtime/lifecycle/start_model.go applies FallbackModel only when the requested model is missing from the executor's advertised ACP catalog. No path applies it when an advertised model later rejects a prompt with a rate limit or quota error.
Rate limits are classified, but not as "this model is unusable".routingerr defines quota_limited / rate_limited (apps/backend/internal/agent/runtime/routingerr/routingerr.go:38), and IsAvailabilityCode deliberately excludes transient rate limiting and quota from the "model unavailable, change the model" guidance.
On a task session, a quota failure currently ends in a manual card. A reset time is surfaced but "never schedules a Kanban retry" (docs/specs/agents/requirements/agent-stall-recovery.md:124). Transient codes get a bounded same-model retry ladder (apps/backend/internal/orchestrator/event_handlers_transient.go); a hard quota failure does not.
The machinery that would do this exists but is gated off. Provider/model-scoped degraded health with a retry time (apps/backend/internal/office/repository/sqlite/provider_health.go), reset waits (routingpolicy.ResetWaitPolicy, decision waiting_for_reset in apps/backend/internal/agent/runtime/routingpolicy/policy.go) and quota circuits (apps/backend/internal/agent/runtime/dynamic/circuit.go) all belong to dynamic agent routing, which is false in every shipped profile (profiles.yaml:75) and marked experimental/high risk (apps/backend/internal/runtimeflags/registry.go:147).
Net effect: one rate-limited provider stalls the task on a manual recovery card even when the profile already has a fallback model configured and the provider reported a reset time.
Proposed solution
For a model that stops on a rate limit or quota:
Mark that provider+model failed/unavailable durably, with the provider reset time when known, so later work that would use the same provider+model skips it until the reset passes or a bounded maximum wait expires.
If the agent profile has a fallback model configured and the executor advertises it, fall back immediately: continue the session on the fallback model without human action, and make the switch visible in the session.
If no fallback is configured, do not switch models. Hold the model unavailable until the end of the bounded wait derived from the reset time, then resume or retry the same model in the background.
Affected area
Agent lifecycle
Who needs this?
Individual developer
Target workflow
Create a task on the Kanban board and run it with a concrete agent profile that has a model (optionally a fallback model) configured.
The provider rate-limits or exhausts quota for that model mid-turn.
Today: the run ends, shows a recovery card naming the model and reset time, and waits for the user.
Wanted: the run continues on the fallback model, or resumes by itself after the known reset when no fallback exists.
Office autonomous sessions were not assessed by the reporter.
Alternatives considered
Manual recovery card (today). Works, but needs a human and never uses the profile's fallback model.
Dynamic agent routing (already implemented, flag-gated off). Provides cross-candidate fallback, reset waits, and provider/model circuits, but needs the experimental features.dynamicAgentRouting flag and a dynamic profile whose fallback is a different concrete profile, not the current profile's fallback model.
AutoFallback on the profile. Session-start only, and it explicitly ignores the configured FallbackModel, so it cannot recover a rate-limited running model.
Acceptance criteria
Given a running session whose profile model returns a rate-limit/quota failure, when the profile has an advertised fallback model, then Kandev switches that session to the fallback model and continues without human action, and the switch is recorded in the session.
Given the same failure with no fallback configured, when the provider reports a reset inside the bounded maximum wait, then the model is marked unavailable until that reset and the session resumes or retries the same model in the background without a manual resume.
Given an unknown reset, or one beyond the bounded maximum wait, then the run stops with the existing recovery surface and the model stays marked unavailable.
The unavailable mark is scoped to provider+model, is used by later sessions on any profile that would use it, and clears when the reset passes.
A strict profile (Require exact model on) never silently switches models; it fails visibly instead.
Switching models stays visible and does not break the existing "one durable warning per selection decision" rule.
No provider raw text beyond the existing sanitized, bounded diagnostic plus the reset time is persisted.
Behavior when no rate limit occurs is unchanged.
Risks and constraints
Contradicts two current documented rules: the Kanban "a reset time never schedules a Kanban retry" rule (agent-stall-recovery.md:124) and the "No silent model fallback" contract (docs/specs/agents/requirements/no-silent-model-fallback.md). A fallback must stay visible and must not weaken Require exact model, so this needs a deliberate spec change, not only an implementation tweak.
Should reuse the existing classification, health, and reset-wait machinery rather than adding a parallel mechanism.
Provider reset times can be missing, wrong, or far in the future, so a maximum wait bound is required (the existing routing policy caps reset waits at 7 days, far too long for interactive work).
Office autonomous runs were not evaluated and may need a separate decision.
Problem
When the model bound to an agent profile stops on a provider rate limit or quota (for example a Codex/OpenCode usage-limit error), Kandev treats it as a session failure that needs a human, not as the model going away. The profile's configured fallback model is never used.
Current behavior, verified in this repository:
StartModelPolicy(FallbackModel,AutoFallback,RequireExactModel) inapps/backend/internal/agent/runtime/lifecycle/start_model.goappliesFallbackModelonly when the requested model is missing from the executor's advertised ACP catalog. No path applies it when an advertised model later rejects a prompt with a rate limit or quota error.routingerrdefinesquota_limited/rate_limited(apps/backend/internal/agent/runtime/routingerr/routingerr.go:38), andIsAvailabilityCodedeliberately excludes transient rate limiting and quota from the "model unavailable, change the model" guidance.docs/specs/agents/requirements/agent-stall-recovery.md:124). Transient codes get a bounded same-model retry ladder (apps/backend/internal/orchestrator/event_handlers_transient.go); a hard quota failure does not.apps/backend/internal/office/repository/sqlite/provider_health.go), reset waits (routingpolicy.ResetWaitPolicy, decisionwaiting_for_resetinapps/backend/internal/agent/runtime/routingpolicy/policy.go) and quota circuits (apps/backend/internal/agent/runtime/dynamic/circuit.go) all belong to dynamic agent routing, which isfalsein every shipped profile (profiles.yaml:75) and marked experimental/high risk (apps/backend/internal/runtimeflags/registry.go:147).Net effect: one rate-limited provider stalls the task on a manual recovery card even when the profile already has a fallback model configured and the provider reported a reset time.
Proposed solution
For a model that stops on a rate limit or quota:
Affected area
Agent lifecycle
Who needs this?
Individual developer
Target workflow
Office autonomous sessions were not assessed by the reporter.
Alternatives considered
features.dynamicAgentRoutingflag and a dynamic profile whose fallback is a different concrete profile, not the current profile's fallback model.AutoFallbackon the profile. Session-start only, and it explicitly ignores the configuredFallbackModel, so it cannot recover a rate-limited running model.Acceptance criteria
Require exact modelon) never silently switches models; it fails visibly instead.Risks and constraints
agent-stall-recovery.md:124) and the "No silent model fallback" contract (docs/specs/agents/requirements/no-silent-model-fallback.md). A fallback must stay visible and must not weakenRequire exact model, so this needs a deliberate spec change, not only an implementation tweak.References
docs/specs/agents/requirements/agent-stall-recovery.md,docs/specs/agents/requirements/no-silent-model-fallback.md,docs/specs/agents/requirements/dynamic-agent-routing.mdapps/backend/internal/agent/runtime/lifecycle/start_model.go,apps/backend/internal/agent/runtime/routingerr/routingerr.go:38,apps/backend/internal/agent/runtime/routingpolicy/policy.go,apps/backend/internal/agent/runtime/dynamic/circuit.go,apps/backend/internal/office/repository/sqlite/provider_health.go,profiles.yaml:75Before submitting