Pick the preset matching your GPU VRAM budget:
llama-server --models-dir D:\AI\LLM\gguf --models-preset presets\models_16GB_VRAM.ini --models-max 1llama-server --models-dir D:\AI\LLM\gguf --models-preset presets\models_24GB_VRAM.ini --models-max 1Dual-GPU (16 GB + 8 GB across two cards) - set the device order before launching so the
per-device fit-target values line up with the physical cards:
$env:CUDA_DEVICE_ORDER = "PCI_BUS_ID" # CUDA0 = 8 GB GPU, CUDA1 = 16 GB GPU
llama-server --models-dir D:\AI\LLM\gguf --models-preset presets\models_16GB_8GB_VRAM.ini| Flag | Purpose |
|---|---|
--models-dir |
Directory containing GGUF files (router mode source #1) |
--models-preset |
INI file with model configs (router mode source #2) |
Tip
main-gpu, models-max, split-mode, tensor-split, and threads are handled by the
preset's [*] global section. --host, --port, and --models-dir must stay on the CLI
— they are parent-server settings that the server manages internally and cannot be set via preset.
Note
The presets models_16GB_VRAM.ini, models_24GB_VRAM.ini, and models_16GB_8GB_VRAM.ini are each tuned for its VRAM budget (context size, KV quantisation, and MoE offload differ). Copy one as a starting point for other hardware. Only models_16GB_8GB_VRAM.ini pins GPUs — the other two leave split-mode/tensor-split unset, so on a multi-GPU host they spread across every visible CUDA device and may exceed the budget named in the file. Pin with CUDA_VISIBLE_DEVICES before launching; its indices follow CUDA_DEVICE_ORDER, which defaults to FASTEST_FIRST and does not match nvidia-smi ordering, so a GPU UUID is the unambiguous choice. Every entry uses load-mode = dio except Qwen3.8-Flash-Next, which needs mmap so that its 26.8 GiB n-gram embedding table can be read on demand — do not normalise that one away.
Important
models_16GB_8GB_VRAM.ini (dual-GPU: one GPU with 16 GB VRAM + one with 8 GB VRAM).
It uses split-mode = layer (pipeline parallel - the recommended mode for consumer GPUs
on PCIe without NVLink). The [*] section sets tensor-split = 1,2 to weight the 16 GB
card twice as heavily as the 8 GB card; the two Qwen3.8-27B entries override it to
1,3, which buys prompt processing at the cost of 16 GB-card VRAM per context token.
main-gpu = 1 puts scratch buffers and intermediate results
on the 16 GB card. models-max = 1 limits to one loaded model at a time. Every entry
except Qwen3.8-Flash-Next sets fit = off with a fixed ctx-size and
n-gpu-layers = -1 baked in for deterministic launches.
- Device order matters.
tensor-split,main-gpu, and--deviceall follow llama.cpp's CUDA order (shown byllama-server --list-devices), not thenvidia-smiorder. SetCUDA_DEVICE_ORDER=PCI_BUS_ID(as above) soCUDA0is the 8 GB card andCUDA1is the 16 GB card, then verify once with--list-devices. - Vision entries split two ways. The two
Qwen3.8-27Bentries setmmproj-device = CUDA0, putting the projector on the card that is not saturated — 4.5x faster image prefill than running CLIP on the CPU. Thegemma-4,Muse-Glimmer-30BandQwen3.8-Flash-Nextentries setno-mmproj-offload = trueinstead, which runs CLIP on the CPU; forQwen3.8-Flash-Nextthat is required rather than a choice, because on afit = onentry the projector loads after the expert split is already committed. - Never set both keys on one entry.
mmproj-deviceandno-mmproj-offloadwrite the same field and preset keys are emitted in unordered-map order, so the winner is whichever lands last. Use one. And on a VRAM-saturated card, letting the projector onto the GPU can OOM the CLIP warmup buffer silently and only fail at image-generation time, so verify a change with a real image request rather than a clean startup.
Each [section] is a model. Keys are llama-server flags without -- for example:
[model-name]
model = /path/to/file.gguf
n-gpu-layers = -1
ctx-size = 262144
parallel = 2The section header (e.g. [gemma-4-31B-it.IQ4_XS.gguf]) is the model name clients pass in the OpenAI-compatible "model" field.
Tip
See llama-server --help for all flags.
Important
All Qwen3.8-*, Ternary-Bonsai-27B and Bonsai-27B entries set
chat-template-file = vendor\Qwen-Fixed-Chat-Templates\chat_template.jinja,
overriding the buggy template embedded in the GGUF. The vendored template is a
single unified file that handles Qwen 3.5, 3.6 and 3.8 variants; both Bonsai models are
Qwen3.6-27B derivatives with a byte-identical tokenizer and embedded template. The path is
repo-relative, so launch llama-server from the repository root (as the examples above do).
If you cloned without --recurse-submodules, run git submodule update --init
first — otherwise startup fails with a missing-file error.
The Bonsai entries additionally set reasoning-effort = medium, the Qwen3.8-*
entries xhigh. medium is the one level that injects no instruction text into the system
prompt, and Qwen 3.6 — which both Bonsai models derive from — has no trained notion of the
concept; Qwen 3.8 is trained on it and
xhigh is what Qwen's own template defaults to. Both are pinned rather than left unset because
the vendored template's own default has moved between its releases, so an unpinned entry would
silently change reasoning level at a template bump. Clients can still override per request via
the OpenAI reasoning_effort field.
All gemma-4-* entries set chat-template-file = vendor\llama.cpp\models\templates\google-gemma-4-31B-it.jinja —
the official Google template bundled with llama.cpp itself, kept in lock-step with
its built-in Gemma 4 parser. The same repo-root launch caveat applies.
Qwen3-Coder-Next entries continue to use their GGUF-embedded template.