- Engine projections never ran:
QwenBlock.load_weightsreadin_featuresfrom the scales tensor (number of groups, e.g. 7), so the INT4 GEMV guard skipped every projection and left garbage. It is nowpacked.size(1) * 8, with shape and dtype checks. - Qwen2 architecture: the block now implements grouped-query attention (2 KV heads; query head h reads KV head h / 7; k/v are 896→128) and the q/k/v biases. RoPE and the KV-cache use the correct row strides.
- Engine safety: checks for tokens per forward, KV-cache overflow and rollback range. CUDA errors surface as Python exceptions. The block owns the weight tensors it points to and cannot be copied.
- Flash-Decoding race: the reduced score shared
warp_sums[0]with warp 0's partial sum (a write-after-read hazard across warps). It now has its own shared variable. - RoPE precision: under
-use_fast_math,powf/sinf/cosfcompiled to approximate instructions.- Qwen2.5's large q/k biases amplify the resulting ~1e-5 rad angle error: the engine drifted from Hugging Face by ~1.6% after layer 0 and ~12% in the logits.
- The angle now follows Hugging Face's fp32 computation, and sin/cos are evaluated in double precision.
- Kernels: the SwiGLU literal
1.0was double precision, forcing FP64 math. The INT4 GEMV guard now requirescols % 32. - Level 5 benchmark: a 1-D input was interpreted as 896 tokens, and the synthetic scales had the wrong shape. The benchmark now measures real weights at fixed context lengths.
- Build:
- CMake links
CUDA::cudart, so the host.cppfiles findcuda_runtime.h, and builds for the native architecture. - Docker uses CUDA 13.0, which matches
torch 2.13.0+cu130; sm_120 needs ≥ 12.8. - The project is no longer installed as a package (
[tool.uv] package = false). build.shuses one CUDA toolkit for CMake and the extensions ($CUDA_HOMEor the firstnvccon PATH) and checks its major version against torch. CMake would otherwise pick an nvcc next to the C++ compiler, such as Ubuntu's CUDA 12.0 in/usr/bin.
- CMake links
- HumanEval evaluation:
- The bigcode-evaluation-harness arguments are complete.
- INT4 prefill now runs on the INT4 GEMM.
- Unit-test execution is serialized, because filelock 3.32 raises on concurrent forks from threads.
- The dataset is loaded as
openai/openai_humaneval. Currenthuggingface_hubrejects bigcode's pre-namespace idopenai_humaneval, and bigcode swallowed the error.
- Measurements:
- Level 5 reports resident VRAM: torch tensors plus the C++
cudaMallocbuffers, which torch statistics do not see. The load-time INT4 quantization peak is reported separately. - Level 6 speeds are decode-only: the prompt prefill is identical in both modes and is reported separately.
- Level 5 reports resident VRAM: torch tensors plus the C++
- Full 24-layer engine (
scripts/engine.py): chunked prefill, greedy generation and a config check against the kernels' hard-coded choices. - Level 6 end to end: Prompt Lookup Decoding on the full model, with rollback in every layer, acceptance and speedup metrics, and a lossless check against greedy decoding.
- C++ harness:
- Random data and CPU references (double accumulation). INT4 kernels are checked against the dequantized weights.
- NaN-prefilled outputs,
CUDA_CHECK, and validation of every kernel and shape, including Split-K. - L2-hot and L2-cold timings with p50/p90/p99, and JSON output.
- Engine tests use a float64 Hugging Face reference with a tolerance calibrated to Hugging Face's own fp32 error, plus a small model with large q/k biases that reproduces Qwen2.5's conditioning without a download.
- Tests (
pytest): CPU tests for packing, thenn.Lineardrop-in, the roofline and the kernel emulation. GPU tests for the kernels and for the engine against Hugging Face Qwen2 (block, logits, greedy generation, speculative decoding). pytestprints the GPU and extension status in its header, and GPU tests are skipped with the real import error.scripts/roofline.py(bandwidth ceilings),scripts/results_io.py(results with environment and git commit),scripts/report.py(README tables),scripts/llama_cpp_benchmark.py(F16 / Q4_0 / Q4_K_M via llama-bench),build.sh.
- One model everywhere:
Qwen/Qwen2.5-Coder-0.5B(base). Level 0 uses a raw completion prompt instead of a chat template. - Flash-Decoding validation:
- Standard-normal inputs, tolerance 1e-4.
- Dominant-key cases, and a check that the data would catch a kernel ignoring the scores.
- The benchmark runs at the model's head_dim (64).
- Python ≥ 3.12;
pytestdev dependency;ninjafor faster extension builds.