Repository navigation
feat(inference): --quantization and --kv_cache_dtype; bump interp-engine to 1.8.0 - #241
Merged
Merged
Conversation
…ine to 1.8.0 start.py grows two flags that reach interp_engine.load_model by the same names (MODEL_QUANTIZATION / KV_CACHE_DTYPE override them, as env does for every flag): --quantization fp8|bnb-4bit|bnb-8bit quantizes a wider checkpoint as it loads, embeddings kept at --model_dtype, and --kv_cache_dtype fp8 halves vLLM's paged cache. Only set values are passed, so a pod that sets neither makes the same load_model call as before; a venv on an older engine is refused with the fix rather than a TypeError from a constructor. The KV cache dtype also feeds the serving limits, so an fp8 cache admits the concurrency it can hold. FP8 on load is what fits Llama-3.3-70B-Instruct on one 96 GB card (local_scripts/pods.yaml inference-llama-3.3-70b-it-fp8). 1.8.0 is the release that adds quantization= and kv_cache_dtype= to load_model. graph and nla move to the same pin for lockstep only: graph builds an EagerModel directly and nla already passes its own NLA_FP8_VERBALIZER / NLA_KV_CACHE_DTYPE into VLLMModel, so neither gains a flag here.
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
What
start.pygrows two flags that reachinterp_engine.load_modelby the same names (envMODEL_QUANTIZATION/KV_CACHE_DTYPEoverride them, as env does for every flag):--quantization fp8|bnb-4bit|bnb-8bitquantizes a wider checkpoint as it loads. Linear layers narrow; embeddings stay at--model_dtype.fp8is vLLM-only,bnb-4bitworks on both backends,bnb-8biton eager; the engine refuses a scheme the resolved backend cannot apply.--kv_cache_dtype fp8halves vLLM's paged KV cache. The serving limits price the cache at its own dtype, so an fp8 cache admits the concurrency it can hold.Only set values are passed, so a pod that sets neither makes the same
load_modelcall as before. A venv on an engine older than 1.8 is refused at startup with the fix (uv sync) rather than dying in a constructor with an opaqueTypeError. Both appear in the config summary and the ready banner.FP8 on load is what fits Llama-3.3-70B-Instruct on one 96 GB RTX PRO 6000: ~67.7 GiB of weights (fp8 linear + 3.9 GiB of bf16 embeddings) against 131 GiB stored, per the engine's
gpu-sizer. The pod config lives inlocal_scripts/pods.yaml(inference-llama-3.3-70b-it-fp8), which is gitignored.interp-engine 1.8.0
1.8.0 (decoderesearch/interp-engine#18) is the release that adds
quantization=andkv_cache_dtype=toload_model. graph and nla move to the same pin for lockstep only: graph builds anEagerModeldirectly and its attribution code reads raw weights, and nla already passes its ownNLA_FP8_VERBALIZER/NLA_KV_CACHE_DTYPEintoVLLMModel, so neither gains a flag here.Tests
tests/unit/test_load_precision.py(new, 7 tests): flag to env toload_modelhandoff, the old-engine refusal, KV pricing,args.pydefaults.check_lint_config_parity.py,check_no_local_path_deps.pypass.