Skip to content

fix: let a static pod declare the layer a readout axis reads - #227

Merged
hijohnnylin merged 1 commit into
mainfrom
static-points-extra
Aug 29, 2026
Merged

hijohnnylin merged 1 commit into
mainfrom
static-points-extra

Conversation

@hijohnnylin

Copy link
Copy Markdown
Owner

What broke

The 70B pod 500s on any assistant-axis request:

ValueError: Capture during generation asked for resid_post.40, which this engine
did not declare. It was built with backend='vllm-static', whose tap set is fixed
when the CUDA graphs are recorded, so no new tap can be installed on the running
model. Declared: resid_post.50

STATIC_POINTS=sae declares exactly the sites the loaded SAEs read, and
resid-post-gf is a single SAE at layer 50. The lu_assistant-axis row for
llama3.3-70b-it reads layer 40.

This only became possible in #226. A shipped asset was on the pod's disk before
the graphs were recorded, so which layers it would be asked for was a deploy-time
fact; an axis that arrives with the request makes it a per-request one, and
nothing on the pod can derive it at startup.

The 8B pod is unaffected — it runs auto, which already covers its axes at
layers 13 through 29.

Why not an existing value

Neither alternative fits one extra layer:

  • sae+auto covers it, but declares every layer. pods.yaml records this exact
    pod as "could take the flip and is refused it on memory"; one layer costs
    1/80th of it.
  • An explicit JSON list is documented as the only value that declares no writes,
    so it would trade every steer for the readout.

The change

STATIC_POINTS_EXTRA adds sites beside a set the server resolves for itself
(sae / sae+auto), as a JSON list in the same spelling — ["resid_post.40"].
Declared as reads and writes, for the reason auto implies its writes: a
capture site that cannot be written is a readout that works and then refuses
every steer at the same layer, and steering on a persona direction at the layer
it was fitted at is a thing this server is asked to do. Deduplicated, so naming a
site an SAE already covers is free. Refused rather than ignored on any other
STATIC_POINTS value, where it would be a no-op that reads, in a deploy config,
as though a site had been added.

The refusal moves earlier and gets specific. assert_residual_available asks
whether a point is declared, which is right for an endpoint that reads every
layer or none — an axis reads exactly one, so resid_post looked present and
only the address was missing. assert_capture_layers_declared checks by address
before anything is rendered or generated, and returns a 400 naming the axis, the
missing layer and the variable to set, instead of a traceback out of the engine.

The pod's own STATIC_POINTS_EXTRA: '["resid_post.40"]' lives in gitignored
local_scripts/pods.yaml, alongside a sae_memory.py change so the freeze-buffer
estimate counts the two extra buffers (32 KiB/token on d_model 8192).

Test plan

  • ruff check + ruff format --check
  • Unit tests for the parse (both address spellings, malformed input, unset),
    the union (reads and writes, dedup), and the 400 (by layer, not by point
    name; a hooked pod still captures anywhere)
  • Full inference unit suite + pyright, run against a byte-identical tree in a
    fully-installed venv — this checkout's apps/inference/.venv is a partial
    install with no torch, so pytest cannot load conftest.py here
  • Not yet exercised on a real static pod. The layer-40 capture is only
    proven by relaunching inference-llama-3.3-70b-it-static on this build

Made with Cursor

A `vllm-static` pod fixes its tap set when the CUDA graphs are recorded, and
`STATIC_POINTS=sae` declares exactly the sites the loaded SAEs read. That was
enough while an axis was a directory on the pod's disk: the layer was there
before the graphs were, so a mismatch was a deploy-time fact.

An axis is a `Vector` row now and arrives with the request, so the layer it
reads is not knowable at startup and cannot be added to a running pod at all.
The 70B pod declares the layer-50 site its one SAE reads and is asked for layer
40 by `lu_assistant-axis`, which fails from inside `generate_steered` -- a 500
with a traceback, after the prompt was rendered.

Neither existing value fits one extra layer. `auto` declares every layer, which
is the flip `inference-llama-3.3-70b-it-static` is refused on memory, and an
explicit JSON list is the one value that declares no writes, so it would trade
every steer for the readout. `STATIC_POINTS_EXTRA` adds sites beside a resolved
`sae` / `sae+auto` set, in the same spelling, as reads AND writes -- both, so a
persona direction can still be steered at the layer it was fitted at. It is
refused rather than ignored on any other value, which would have nothing to add
to, since a no-op there would read in a deploy config as though a site had been
added.

The refusal moves too. `assert_residual_available` asks whether a point is
declared, which is the right question for an endpoint reading every layer or
none; an axis reads exactly one, so `resid_post` looked present and only the
address was missing. `assert_capture_layers_declared` checks by address before
anything is generated, and names the axis, the layer and the variable to set.

Co-authored-by: Cursor <cursoragent@cursor.com>
Comment thread apps/inference/neuronpedia_inference/endpoints/steer/completion_chat.py Dismissed
@hijohnnylin
hijohnnylin merged commit 9b5167b into main Aug 29, 2026
16 of 17 checks passed
@hijohnnylin
hijohnnylin deleted the static-points-extra branch August 29, 2026 01:54
hijohnnylin added a commit that referenced this pull request Aug 29, 2026
… a 400 (#228)

* feat: report the declared tap set in both startup banners

Which sites a pod declared decided whether a request would be served, and was
not written down anywhere a human reads at startup. It took a 500 mid-stream to
find out that a pod's tap set and its axis layers disagreed.

Both banners now say so, and they answer different questions. The config banner
prints before the SAEs load, so `sae` has no address list yet and says that
rather than an empty set that would read as "nothing declared". The ready banner
is the same question after resolution and after the extras merge in, read off the
loaded model.

Collapsed to ranges because the honest list is unreadable: `auto` on this 70B pod
declares 80 addresses per direction, which is `resid_post.0-79` here. One token
per point name, layers grouped inside it, so `mlp_out_post.0-25 resid_mid.0-25`
is a whole gemmascope set.

Co-authored-by: Cursor <cursoragent@cursor.com>

* fix: refuse a steer at an undeclared layer with a 400, not a 500 mid-stream

The read-side twin of this landed in #227 and left the write side asymmetric: an
axis reading an undeclared layer got a 400 naming the layer, while a steer
WRITING to one still died inside generate_steered, several hundred tokens into a
StreamingResponse that had already returned 200.

`assert_steering_available` did not catch it because it asks whether the pod
declared any write site at all, which is the generation-only case. What bites is
writes declared, at other layers: a projection cap writes wherever the vector it
caps was fitted, so on the assistant-axis endpoint the pod declared 40 for the
readout and was asked to write 32. An uploaded vector can name any layer, so the
set is not one a pod can enumerate at startup.

Both checks now match sites the way the engine does, through the
`resid_pre[L]` == `resid_post[L-1]` alias, so a pod that declared one spelling is
not refused for the other. They lean permissive on anything they do not
recognize: an unrecognized site is one the engine refuses for itself, which is
the failure they replace, where a false refusal would break a request the pod
could serve.

`steer_write_layers` shares `steer_layer_for_hook` with the spec that does the
writing, since computing the layer twice is how a check passes for one layer and
the engine fails at another. It runs only on a pod whose taps are fixed --
resolving those layers reads each feature's hook out of the SAE manager, and a
check that cannot refuse anything should not be able to raise on the way to
saying so.

Co-authored-by: Cursor <cursoragent@cursor.com>

---------

Co-authored-by: Cursor <cursoragent@cursor.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants