Skip to content

chore: bump interp-engine to 1.6.0 in inference, graph and nla - #233

Merged
hijohnnylin merged 1 commit into
mainfrom
interp-engine-1.6
Sep 5, 2026
Merged

hijohnnylin merged 1 commit into
mainfrom
interp-engine-1.6

Conversation

@hijohnnylin

Copy link
Copy Markdown
Owner

Moves the three apps that import the engine — inference ([vllm,quant]), nla ([vllm]) and graph (base) — from interp-engine 1.3.3 to 1.6.0. autointerp and sparsity never imported it, so they are untouched.

The pin is exact by design, so it only moves deliberately. Taking 1.6.0 for its speed work and for half the static-buffer VRAM on vllm-static, plus the Gemma 4 feed-forward, router and value point fixes that landed in 1.4.0.

No call sites move

Between 1.3.3 and 1.6.0 the engine's public surface is additive. __init__.py only gains HubKernelUnsupported and deepgemm_fallback_kwargs, and the one changed public signature — kv_heads_for_layer — is not called from this repo. The new interp_engine.memory sizer (1.5.0) is unused here so far. So this is a pin bump, a relock, and the version named in AGENTS.md; no Python changed.

Worth knowing at deploy time

The engine owns the vLLM pin, so this drags vLLM 0.27.1 → 0.28.0 through inference's and nla's locks. That is the largest real change in the diff and the thing to watch when the next pod comes up.

Verification

Run locally on an RTX 5090 with HF_TOKEN set, so the gated-model tests really ran rather than skipping themselves.

App Tests Pyright
inference 691 passed, 31 skipped, 3 xfailed 0 errors
graph 66 passed, 10 skipped 4 pre-existing
nla 27 passed 0 errors
autointerp 21 passed 0 errors
sparsity 3 passed 0 errors

The inference run exercised real gpt-oss-20b, llama3.1-8b-it and gemma-2-2b, plus the vLLM batch-layout paths on both v1 and v2. make python-lint and make openapi-check pass, and the local-path-dependency gate confirms nothing leaked into the locks.

Every skip is opt-in rather than an environment gap: inference's 31 want NP_RUN_MANUAL_TESTS=1 (16 GB of gated weights) and graph's 10 want RUN_CRM_PARITY=1 plus the crm extra.

Graph's 4 pyright errors are pre-existing transformers typing in steer_generation.py and its test. I reinstalled 1.3.3 and got the identical 4, so they predate this change and are left alone for a separate fix.

The pin is exact, so it only moves deliberately. Taking 1.6.0 for its speed
work and for half the static-buffer VRAM on vllm-static, plus the Gemma 4
feed-forward, router and value point fixes that landed in 1.4.0.

No call sites move. Between 1.3.3 and 1.6.0 the engine's public surface is
additive -- __init__ only gains HubKernelUnsupported and
deepgemm_fallback_kwargs -- and the one changed public signature,
kv_heads_for_layer, is not called from this repo. The new interp_engine.memory
sizer is unused here so far.

Worth knowing at deploy time: the engine owns the vLLM pin, so this drags vLLM
0.27.1 -> 0.28.0 through inference's and nla's locks. That is the largest real
change in the diff.

Verified on an RTX 5090 with HF_TOKEN set: pyright clean for inference and nla,
and 691 passed / 31 skipped / 3 xfailed in inference including the gpt-oss-20b,
llama3.1-8b-it and gemma-2-2b models and the vLLM batch-layout paths, 66 in
graph, 27 in nla. Graph's 4 pyright errors are pre-existing transformers typing
in steer_generation.py -- identical with 1.3.3 installed -- and are left alone.

Co-authored-by: Cursor <cursoragent@cursor.com>
@hijohnnylin
hijohnnylin merged commit 8cd2b94 into main Sep 5, 2026
20 of 21 checks passed
@hijohnnylin
hijohnnylin deleted the interp-engine-1.6 branch September 5, 2026 00:29
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant