View from downstream: clients can't discover the served context length, and it fails silently #1786
behrnt-slatgng
started this conversation in
General
Replies: 0 comments
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Uh oh!
There was an error while loading. Please reload this page.
Disclosure up front: I build Grunz, a hosted chat + coding agent running abliterated open weights. I'm a consumer of endpoints rather than an operator of one, so this is a view from downstream. Nothing to sell here — it's a report I think is more useful to operators than to users.
The gap between the model card and the served window
The most common surprise when moving a workload between hosted endpoints is that the advertised context window on a model card has almost no relationship to the window you actually receive.
Across the endpoints I've tested, most serve around 32K regardless of what the card claims. A handful reach 256K. I have not found one actually serving the 1M figures that appear in release posts.
From the operator side this is entirely reasonable — max model length is a KV-cache budget decision traded off against concurrency, not a marketing decision. The problem is that it's almost never documented and there's no standard way for a client to discover it. So what clients actually do is:
The second one is genuinely dangerous, because a silently truncated run doesn't error. It just quietly gets worse.
Why this bites roleplay and agent workloads specifically
For short chat, 32K is fine and nobody notices. For long-context work it's the binding constraint, and it fails in a way that doesn't look like a context failure.
In roleplay terms: the character starts forgetting established facts, drifting out of character, contradicting itself. Users blame the model or their own card, and the standard advice ("trim the lorebook", "shorten the character definition") treats the symptom. In agent terms: the run hits the ceiling, compaction fires, the goal gets summarised into vagueness, and the model reads back its own compacted notes and restarts the plan.
Either way it reads from outside as "this model is bad at long contexts." It's the serving config.
A second finding that may matter to you as an engine
Abliterated models degrade instruction-following and output-format adherence before they degrade knowledge. The model still knows the material; it gets worse at respecting the instruct template, prefill, stop sequences, and any structured-output contract.
That has a direct implication for anyone serving these: grammar-constrained or structured decoding is doing more load-bearing work on an abliterated model than on a clean one. If you're benchmarking engine features against base models, the abliterated case may behave quite differently, and given what Aphrodite is mostly used for, that's probably the case that matters.
The ask
Is there — or would there be appetite for — a discoverable way for clients to read the effective served context length, through model metadata rather than reverse-engineering it from error strings? I may have missed an existing mechanism, in which case I'd like to be pointed at it.
And if anyone has measured how abliterated models behave under constrained decoding compared to their base, I'd genuinely like to read it. My evidence is anecdotal and I'd rather cite a result.
All reactions