Skip to content

fix: make a static pod's tap set visible, and refuse a bad steer with a 400 - #228

Merged
hijohnnylin merged 2 commits into
mainfrom
static-points-log
Aug 29, 2026
Merged

hijohnnylin merged 2 commits into
mainfrom
static-points-log

Conversation

@hijohnnylin

Copy link
Copy Markdown
Owner

Follow-up to #227. Two changes, both about the same thing: which sites a pod
declared decides whether a request is served, and that was neither visible at
startup nor reported cleanly when it went wrong.

1. The declared tap set is in both startup banners

Config banner, before the SAEs load:

Initialized Config with:
  model_id: meta-llama/Llama-3.3-70B-Instruct
  ...
  static_points_summary: sae (resolved once the SAEs load) + resid_post.40

Ready banner, read off the loaded model after resolution:

==== LOADING COMPLETE - SERVING ON http://localhost:5002 ====
  ...
  static taps: read resid_post.0-79 | write resid_post.0-79
  startup took 6m 12s

Two lines because they answer two questions. sae has no address list until the
SAEs load, so the first says that rather than printing an empty set that would
read as "nothing declared"; the second is the truth after the extras merge in. A
pod missing a layer its axes need is now visible at startup rather than on the
first request that needs it.

Ranges because the literal list is unreadable — auto on the 70B pod is 80
addresses per direction. One token per point name with layers grouped inside it,
so a gemmascope set reads mlp_out_post.0-25 resid_mid.0-25.

2. A steer at an undeclared layer is a 400, not a 500 mid-stream

#227 fixed the read side and left the write side asymmetric. An axis reading an
undeclared layer got a 400 naming the layer; a steer writing to one still died
inside generate_steered, well into a StreamingResponse that had already
returned 200:

ValueError: Steered generation asked for resid_post.32, which this engine did not
declare. ... Declared: resid_post.40, resid_post.50

assert_steering_available misses this because it asks whether the pod declared
any write site, which is the generation-only case. What bites is writes
declared, at other layers: a projection cap writes wherever the vector it caps
was fitted, so the pod declared 40 for the readout and was asked to write 32. An
uploaded vector can name any layer, so this is not a set a pod can enumerate.

Both checks now match sites the way the engine does, through the
resid_pre[L] == resid_post[L-1] alias, so a pod that declared one spelling is
not refused for the other — the read check in #227 compared literal names and was
stricter than the engine. They lean permissive on anything unrecognized, since
that degrades to the engine's own refusal (the failure they replace) where a
false refusal would break a request the pod could serve.

Deploy note

The pod that prompted #227 no longer uses STATIC_POINTS_EXTRA: capping writes at
layers no list can enumerate, so it runs STATIC_POINTS=auto — priced at
util 0.90 with the prefill chunk pinned to 2048, which leaves KV at ~47,000
tokens. That lives in gitignored local_scripts/pods.yaml, so it is not in this
diff. STATIC_POINTS_EXTRA stays as the lever for a pod that cannot afford auto
and whose axis layers are a closed set; it now has no user.

Test plan

  • ruff check + ruff format --check
  • Ranges: empty, single, full run, run plus outlier, grouping by point name,
    repeats, and a layerless address
  • Config summary for every STATIC_POINTS shape, including a regression test
    that a named mode is not iterated character by character
  • Write check: refusal names the layer and STATIC_POINTS=auto, read and
    write sets asked separately, alias match, and the hooked-pod gate
  • Full inference unit suite (555) + pyright, against a byte-identical tree in
    a fully-installed venv — this checkout's apps/inference/.venv is partial
    and cannot load conftest.py
  • Not exercised on a real static pod. Both refusals and both banner lines
    are only proven by relaunching the 70B pod on this build

Made with Cursor

hijohnnylin and others added 2 commits August 28, 2026 19:26
Which sites a pod declared decided whether a request would be served, and was
not written down anywhere a human reads at startup. It took a 500 mid-stream to
find out that a pod's tap set and its axis layers disagreed.

Both banners now say so, and they answer different questions. The config banner
prints before the SAEs load, so `sae` has no address list yet and says that
rather than an empty set that would read as "nothing declared". The ready banner
is the same question after resolution and after the extras merge in, read off the
loaded model.

Collapsed to ranges because the honest list is unreadable: `auto` on this 70B pod
declares 80 addresses per direction, which is `resid_post.0-79` here. One token
per point name, layers grouped inside it, so `mlp_out_post.0-25 resid_mid.0-25`
is a whole gemmascope set.

Co-authored-by: Cursor <cursoragent@cursor.com>
…stream

The read-side twin of this landed in #227 and left the write side asymmetric: an
axis reading an undeclared layer got a 400 naming the layer, while a steer
WRITING to one still died inside generate_steered, several hundred tokens into a
StreamingResponse that had already returned 200.

`assert_steering_available` did not catch it because it asks whether the pod
declared any write site at all, which is the generation-only case. What bites is
writes declared, at other layers: a projection cap writes wherever the vector it
caps was fitted, so on the assistant-axis endpoint the pod declared 40 for the
readout and was asked to write 32. An uploaded vector can name any layer, so the
set is not one a pod can enumerate at startup.

Both checks now match sites the way the engine does, through the
`resid_pre[L]` == `resid_post[L-1]` alias, so a pod that declared one spelling is
not refused for the other. They lean permissive on anything they do not
recognize: an unrecognized site is one the engine refuses for itself, which is
the failure they replace, where a false refusal would break a request the pod
could serve.

`steer_write_layers` shares `steer_layer_for_hook` with the spec that does the
writing, since computing the layer twice is how a check passes for one layer and
the engine fails at another. It runs only on a pod whose taps are fixed --
resolving those layers reads each feature's hook out of the SAE manager, and a
check that cannot refuse anything should not be able to raise on the way to
saying so.

Co-authored-by: Cursor <cursoragent@cursor.com>
Comment thread apps/inference/neuronpedia_inference/endpoints/steer/completion.py Dismissed
Comment thread apps/inference/neuronpedia_inference/endpoints/steer/completion_chat.py Dismissed
@hijohnnylin
hijohnnylin merged commit 5580619 into main Aug 29, 2026
17 checks passed
@hijohnnylin
hijohnnylin deleted the static-points-log branch August 29, 2026 03:11
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants