-
Notifications
You must be signed in to change notification settings - Fork 391
Pull requests: ikawrakow/ik_llama.cpp
Author
Label
Projects
Milestones
Reviews
Assignee
Sort
Pull requests list
Bucket top_k (CPU): ~3% better TG at 128k context
#2225
opened Aug 1, 2026 by
ikawrakow
Owner
Loading…
window-sized SWA ring KV cache (--swa-compress to opt in)
#2171
opened Jul 23, 2026 by
joaomdsg
Loading…
CUDA sm_60: bound the fp16 flash-attention error on P100 prefill instead of disabling the fast path
#2151
opened Jul 18, 2026 by
mb8565
Contributor
Loading…
ggml : fix build and MoE inference on CPUs without AVX2 (AVX1-only)
#2138
opened Jul 15, 2026 by
neomindryan
Loading…
2 of 4 tasks
llama: disable fused up-gate when an offload backend lacks GGML_OP_FUSED_UP_GATE (fixes silent 3.5x Metal slowdown)
#2133
opened Jul 15, 2026 by
hchengit
Contributor
Loading…
CUDA: update captured graph when a source tensor's shape changes
#2096
opened Jul 8, 2026 by
fibrahimov
Loading…
Be able to force synchronization when copying graph inputs
#2092
opened Jul 6, 2026 by
ikawrakow
Owner
Loading…
Fix misc. expiring logit/sparam bias bugs
#1914
opened Jun 2, 2026 by
dungquixote42
Contributor
Loading…
2 of 4 tasks
server: enable checkpoint reuse for recurrent/hybrid models (qwen3next, Mamba)
#1888
opened May 26, 2026 by
localweights
Loading…
docs: Complete rewrite of build.md – CMake-only, focused on supported backends
#1853
opened May 21, 2026 by
maddes8cht
Loading…
2 of 4 tasks
cuda: add get_rows CUDA kernels for Q4_K, Q5_K, Q6_K
#1830
opened May 18, 2026 by
localweights
Loading…
Slightly expand the usage of VNNI256
#1764
opened May 9, 2026 by
XZiar
Contributor
Loading…
2 of 4 tasks
runtime : add
--run-time-repack auto mode for swap-bound MoE safety
#1738
opened May 4, 2026 by
AndrewMoryakov
Contributor
Loading…
2 of 4 tasks
Change signature of llama_set_draft_input_hidden_state
#1727
opened May 3, 2026 by
ikawrakow
Owner
Loading…
Previous Next
ProTip!
Follow long discussions with comments:>50.