Repository navigation
chore: bump interp-engine to 1.6.0 in inference, graph and nla - #233
Merged
Merged
Conversation
The pin is exact, so it only moves deliberately. Taking 1.6.0 for its speed work and for half the static-buffer VRAM on vllm-static, plus the Gemma 4 feed-forward, router and value point fixes that landed in 1.4.0. No call sites move. Between 1.3.3 and 1.6.0 the engine's public surface is additive -- __init__ only gains HubKernelUnsupported and deepgemm_fallback_kwargs -- and the one changed public signature, kv_heads_for_layer, is not called from this repo. The new interp_engine.memory sizer is unused here so far. Worth knowing at deploy time: the engine owns the vLLM pin, so this drags vLLM 0.27.1 -> 0.28.0 through inference's and nla's locks. That is the largest real change in the diff. Verified on an RTX 5090 with HF_TOKEN set: pyright clean for inference and nla, and 691 passed / 31 skipped / 3 xfailed in inference including the gpt-oss-20b, llama3.1-8b-it and gemma-2-2b models and the vLLM batch-layout paths, 66 in graph, 27 in nla. Graph's 4 pyright errors are pre-existing transformers typing in steer_generation.py -- identical with 1.3.3 installed -- and are left alone. Co-authored-by: Cursor <cursoragent@cursor.com>
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Moves the three apps that import the engine —
inference([vllm,quant]),nla([vllm]) andgraph(base) — frominterp-engine1.3.3 to 1.6.0.autointerpandsparsitynever imported it, so they are untouched.The pin is exact by design, so it only moves deliberately. Taking 1.6.0 for its speed work and for half the static-buffer VRAM on
vllm-static, plus the Gemma 4 feed-forward, router and value point fixes that landed in 1.4.0.No call sites move
Between 1.3.3 and 1.6.0 the engine's public surface is additive.
__init__.pyonly gainsHubKernelUnsupportedanddeepgemm_fallback_kwargs, and the one changed public signature —kv_heads_for_layer— is not called from this repo. The newinterp_engine.memorysizer (1.5.0) is unused here so far. So this is a pin bump, a relock, and the version named inAGENTS.md; no Python changed.Worth knowing at deploy time
The engine owns the vLLM pin, so this drags vLLM 0.27.1 → 0.28.0 through
inference's andnla's locks. That is the largest real change in the diff and the thing to watch when the next pod comes up.Verification
Run locally on an RTX 5090 with
HF_TOKENset, so the gated-model tests really ran rather than skipping themselves.The inference run exercised real gpt-oss-20b, llama3.1-8b-it and gemma-2-2b, plus the vLLM batch-layout paths on both v1 and v2.
make python-lintandmake openapi-checkpass, and the local-path-dependency gate confirms nothing leaked into the locks.Every skip is opt-in rather than an environment gap: inference's 31 want
NP_RUN_MANUAL_TESTS=1(16 GB of gated weights) and graph's 10 wantRUN_CRM_PARITY=1plus thecrmextra.Graph's 4 pyright errors are pre-existing transformers typing in
steer_generation.pyand its test. I reinstalled 1.3.3 and got the identical 4, so they predate this change and are left alone for a separate fix.