Skip to content

Pull requests: ikawrakow/ik_llama.cpp

Author
Filter by author
Loading
Label
Filter by label
Loading
Use alt + click/return to exclude labels
or + click/return for logical OR
Projects
Filter by project
Loading
Milestones
Filter by milestone
Loading
Reviews
Assignee
Filter by who’s assigned
Assigned to nobody Loading
Sort

Pull requests list

Bucket top_k (CPU): ~3% better TG at 128k context
#2225 opened Aug 1, 2026 by ikawrakow Owner Loading…
Fix RPC crash with dflash
#2167 opened Jul 23, 2026 by firecoperana Collaborator Loading…
Be able to force synchronization when copying graph inputs
#2092 opened Jul 6, 2026 by ikawrakow Owner Loading…
add --split-output-tensor / -sot CLI parameter
#2056 opened Jun 29, 2026 by Nexesenex Contributor Draft
2 of 4 tasks
implement perplexity in llama-server
#2011 opened Jun 22, 2026 by magikRUKKOLA Contributor Draft
Fix misc. expiring logit/sparam bias bugs
#1914 opened Jun 2, 2026 by dungquixote42 Contributor Loading…
2 of 4 tasks
Qwen3.5 MTP: extract selected tokens earlier
#1892 opened May 28, 2026 by ikawrakow Owner Loading…
Fix prompt cache viability
#1877 opened May 25, 2026 by zeljkokalezic Loading…
2 of 4 tasks
A GGUF MTP Extract and Merge Tool
#1849 opened May 20, 2026 by FNsi Draft
2 of 4 tasks
Slightly expand the usage of VNNI256
#1764 opened May 9, 2026 by XZiar Contributor Loading…
2 of 4 tasks
runtime : add --run-time-repack auto mode for swap-bound MoE safety
#1738 opened May 4, 2026 by AndrewMoryakov Contributor Loading…
2 of 4 tasks
Change signature of llama_set_draft_input_hidden_state
#1727 opened May 3, 2026 by ikawrakow Owner Loading…
ProTip! Follow long discussions with comments:>50.