Skip to content

fix(models): read a relay's own capacity report and version notation - #1559

Open
pengwei269446668 wants to merge 2 commits into
vastsa:mainfrom
pengwei269446668:fix/relay-model-context
Open

pengwei269446668 wants to merge 2 commits into
vastsa:mainfrom
pengwei269446668:fix/relay-model-context

Conversation

@pengwei269446668

@pengwei269446668 pengwei269446668 commented Oct 11, 2026 •

Copy link
Copy Markdown

Problem

A relay / custom OpenAI-compatible endpoint (e.g. https://sidrune.ai serving deepseek-v4-flash, claude-sonnet-4.5, gemini-3.1-pro) loses the context window its own model list publishes and falls back to the generic 128k / 8192 seed. Older versions (e.g. v0.16.0) resolved these; current main does not.

Root cause: 62cf3c257 fix(models): use models.dev as metadata authority removed the relayChatMetadata fallback in models-dev-catalog.ts, so an unanchored relay id that the catalog does not publish verbatim got nothing. In addition, the service's own capacity report on its model-list row was discarded, and a relay naming a model in the owner's other notation (4.5 vs 4-5) or by the owner's canonical generation id never reached a record.

Fix

  1. Read the endpoint's own capacity report (model-discovery.ts, ipc/provider-ipc.ts). A model-list row that states context_window / context_length / inputTokenLimit / contextWindow (and the matching output keys) keeps that value, and the settings row prefers it over both the catalog and any stored value — PiDeck parity: reported > catalog > nothing. An omission is never guessed.
  2. Notation aliases (models-dev-catalog.ts). A dotted version name (claude-sonnet-4.5) now reads the owner's dashed published record (claude-sonnet-4-5); each alias is checked as a whole published id, next to the existing single-whitelisted-deployment-label strip (-1m, -test, …).
  3. Canonical generation names (models-dev-catalog.ts). A relay naming the owner's canonical_model_id (deepseek-v4.1-flash) is answered by the owner's own live (non-deprecated) record, since the owner publishes that generation only under version aliases.
  4. Owner-record fallback for unanchored relays (models-dev-catalog.ts, pi-model-metadata.ts). When no provider is anchored, only the weight owner's records for the family (claude→anthropic, gpt/o1/o3/o4→openai, gemini→google, grok→xai, mimo→xiaomi, glm→zai, deepseek→deepseek, kimi→moonshotai, minimax→minimax) may answer.

Invariants preserved (approved #1047 matching rules):

  • Reseller copies never decide a window or a capability.
  • The wire id is never rewritten; only whitelisted deployment labels are stripped.
  • Variants (-thinking, dated releases) and generations the catalog does not publish stay unmatched.

Verification

  • models-dev-catalog.test.mjs + model-discovery.test.mjs: 69/69 pass.
  • npx tsc -p tsconfig.json --noEmit: exit 0.
  • Full desktop suite vs the same tree at 9fa3abdee: no new failures (the only deltas are pre-existing, flaky, timeout-bound UI/Markdown tests that pass when run alone).
  • Probes on the fix: deepseek-v4-flash → deepseek ctx=1000000 out=393216, claude-sonnet-4.5 → anthropic as=claude-sonnet-4-5, deepseek-v4.1-flash → deepseek as=deepseek-flash ctx=1000000, mimo-v2.5-pro-1m → xiaomi, some-private-model → UNMATCHED.

Known pre-existing failures (not touched here)

  • provider-endpoint-metadata.test.mjs: on Windows the fixture uses new URL(...).pathname, which yields /E:/… with a leading slash, so readFile fails. Fails on main too.
  • model-binding-catalog-source.test.mjs: catalog drift (1048576 vs 1050000).
  • macOS notarization / plugin / search tests, unrelated to this change.

🤖 Generated with Claude Code

An unidentified relay (a custom OpenAI-compatible endpoint such as a DeepSeek relay) lost its published context window when the models.dev authority landed: without an anchor the lookup dropped to the generic 128k text-only seed even though the owner published the exact id.

Restore the pre-catalog relay rule: the model's own family names the publisher that owns its weights, and that owner's record answers an unanchored row (deepseek-v4-flash reads DeepSeek's 1M window again). Only the owner's exact record answers, one whitelisted deployment label may still be stripped, and the wire id is never rewritten.
A relay that serves an OpenAI-compatible list now states its own capacity,
and its model names may use the owner's notation or canonical generation id.
Three gaps made those rows fall back to the generic 128k/8192 seed:

- A model list's `context_window` / `context_length` / `inputTokenLimit` /
  `contextWindow` (and the matching output keys) was discarded, so a service
  that plainly states 1,000,000 tokens was still enriched from elsewhere.
  Discovery keeps the reported value, and the settings row prefers it over
  both the catalog and any stored value (PiDeck parity: reported > catalog >
  nothing - an omission is never guessed).

- A relay's dotted version name (`claude-sonnet-4.5`) never reached the
  owner's dashed published record (`claude-sonnet-4-5`). Each notation alias
  is now checked as a whole published id, alongside the existing
  single-whitelisted-deployment-label strip.

- A relay naming the owner's canonical generation id (`deepseek-v4.1-flash`)
  found nothing, because the owner publishes that generation only under
  version aliases. The owner's own live (non-deprecated) record for the
  canonical id now answers.

Only the weight owner's records may answer: reseller copies still never
decide a window or a capability, the wire id is never rewritten, and a
variant (`-thinking`, dated releases) or a generation the catalog does not
publish stays unmatched.
@pengwei269446668 pengwei269446668 changed the title fix(models): resolve relay model context from the weight owner's record fix(models): read a relay's own capacity report and version notation Oct 11, 2026

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant