Repository navigation
inference: interp-engine 1.8.0 -> 1.11.0, steer with SteerMethod - #245
Merged
Merged
Conversation
1.11.0 names the steering methods once, as `SteerMethod`, and types `SteerSpec.method` with it. `_feature_to_steerspec` passes the members instead of the strings they equal, so the boundary is typed on both sides. Also in the jump: 1.9/1.10 all-gather q/k/v across tensor-parallel ranks, so the engine no longer refuses attention recompute on `tensor_parallel_size` alone; the contract test that pinned that refusal now pins what replaced it -- a head-sharded capture is named as one when its width is not a whole number of heads. Inference's own pod-level gate for sharded attention is unchanged. 1.10 also pins `vllm==0.28.0` exactly, which is what the lock already resolved, and adds a FlashInfer preflight on vLLM start. Verified: unit suite, eager and vLLM steering integration tests (gpt2 + res-jb). Co-authored-by: Cursor <cursoragent@cursor.com>
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
apps/inferencepininterp-engine[vllm,quant]==1.8.0to==1.11.0; relock (only the engine moved)._feature_to_steerspecpassesSteerMethod.ORTHOGONAL/SteerMethod.ADDITIVEinstead of string literals.TestAttentionRecomputeSharding: since 1.10 the engine all-gathers q/k/v across tensor-parallel ranks and no longer refusesrecompute_attn_from_payloadsontensor_parallel_sizealone. The tests now check that the argument is accepted and that a head-sharded capture is named as a shard. Inference's own pod-level TP gates inengine_adapter.pyare unchanged.Also inherited from 1.9/1.10:
vllm==0.28.0is now an exact pin (already what the lock resolved), and a FlashInfer preflight runs on vLLM start.Verified: ruff, pyright, unit suite (618), and the eager + vLLM steering integration tests (gpt2 + res-jb).
graphandnlastay on 1.8.0; neither uses steering methods.