Skip to content

feat(inference): --quantization and --kv_cache_dtype; bump interp-engine to 1.8.0 - #241

Merged
hijohnnylin merged 1 commit into
mainfrom
inference-quantization-flags
Sep 14, 2026
Merged

hijohnnylin merged 1 commit into
mainfrom
inference-quantization-flags

Conversation

@hijohnnylin

Copy link
Copy Markdown
Owner

What

start.py grows two flags that reach interp_engine.load_model by the same names (env MODEL_QUANTIZATION / KV_CACHE_DTYPE override them, as env does for every flag):

  • --quantization fp8|bnb-4bit|bnb-8bit quantizes a wider checkpoint as it loads. Linear layers narrow; embeddings stay at --model_dtype. fp8 is vLLM-only, bnb-4bit works on both backends, bnb-8bit on eager; the engine refuses a scheme the resolved backend cannot apply.
  • --kv_cache_dtype fp8 halves vLLM's paged KV cache. The serving limits price the cache at its own dtype, so an fp8 cache admits the concurrency it can hold.

Only set values are passed, so a pod that sets neither makes the same load_model call as before. A venv on an engine older than 1.8 is refused at startup with the fix (uv sync) rather than dying in a constructor with an opaque TypeError. Both appear in the config summary and the ready banner.

FP8 on load is what fits Llama-3.3-70B-Instruct on one 96 GB RTX PRO 6000: ~67.7 GiB of weights (fp8 linear + 3.9 GiB of bf16 embeddings) against 131 GiB stored, per the engine's gpu-sizer. The pod config lives in local_scripts/pods.yaml (inference-llama-3.3-70b-it-fp8), which is gitignored.

interp-engine 1.8.0

1.8.0 (decoderesearch/interp-engine#18) is the release that adds quantization= and kv_cache_dtype= to load_model. graph and nla move to the same pin for lockstep only: graph builds an EagerModel directly and its attribution code reads raw weights, and nla already passes its own NLA_FP8_VERBALIZER / NLA_KV_CACHE_DTYPE into VLLMModel, so neither gains a flag here.

Tests

  • tests/unit/test_load_precision.py (new, 7 tests): flag to env to load_model handoff, the old-engine refusal, KV pricing, args.py defaults.
  • inference: 602 unit tests pass; ruff, format, pyright clean.
  • graph: 71 passed / 12 skipped (gpu). nla: 27 passed. Both synced to 1.8.0.
  • check_lint_config_parity.py, check_no_local_path_deps.py pass.

…ine to 1.8.0

start.py grows two flags that reach interp_engine.load_model by the same
names (MODEL_QUANTIZATION / KV_CACHE_DTYPE override them, as env does for
every flag): --quantization fp8|bnb-4bit|bnb-8bit quantizes a wider
checkpoint as it loads, embeddings kept at --model_dtype, and
--kv_cache_dtype fp8 halves vLLM's paged cache. Only set values are
passed, so a pod that sets neither makes the same load_model call as
before; a venv on an older engine is refused with the fix rather than a
TypeError from a constructor. The KV cache dtype also feeds the serving
limits, so an fp8 cache admits the concurrency it can hold.

FP8 on load is what fits Llama-3.3-70B-Instruct on one 96 GB card
(local_scripts/pods.yaml inference-llama-3.3-70b-it-fp8).

1.8.0 is the release that adds quantization= and kv_cache_dtype= to
load_model. graph and nla move to the same pin for lockstep only: graph
builds an EagerModel directly and nla already passes its own
NLA_FP8_VERBALIZER / NLA_KV_CACHE_DTYPE into VLLMModel, so neither gains
a flag here.
@hijohnnylin
hijohnnylin merged commit 599cd3f into main Sep 14, 2026
19 of 20 checks passed
@hijohnnylin
hijohnnylin deleted the inference-quantization-flags branch September 14, 2026 08:39
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant