llama.cpp

History

Yes You Can Have Your Own 50e0ad08fb server: save and clear idle slots on new task (`--clear-idle`) (#20993 ) * server: clear idle slots KV from VRAM (LLAMA_KV_KEEP_ONLY_ACTIVE) * server: move idle slot KV clearing to slot release The save "cost" is now paid by the finishing request. * server: add --kv-clear-idle flag, enable by default * server: skip clearing last idle slot, clear on launch * server: test --no-kv-clear-idle flag * server: simplify on-release clearing loop * server: remove on-release KV clearing, keep launch-only * cont : clean-up * tests: update log strings after --clear-idle rename * tests: use debug tags instead of log message matching * test: fix Windows CI by dropping temp log file unlink --------- Co-authored-by: Georgi Gerganov <ggerganov@gmail.com>		2026-04-03 19:02:27 +02:00
..
batched-bench	common : move up common_init() and fix Windows UTF-8 logs (#21176 )	2026-03-31 12:53:41 +02:00
cli	server: save and clear idle slots on new task (`--clear-idle`) (#20993 )	2026-04-03 19:02:27 +02:00
completion	server: save and clear idle slots on new task (`--clear-idle`) (#20993 )	2026-04-03 19:02:27 +02:00
cvector-generator	common : move up common_init() and fix Windows UTF-8 logs (#21176 )	2026-03-31 12:53:41 +02:00
export-lora	common : move up common_init() and fix Windows UTF-8 logs (#21176 )	2026-03-31 12:53:41 +02:00
fit-params	common : move up common_init() and fix Windows UTF-8 logs (#21176 )	2026-03-31 12:53:41 +02:00
gguf-split	gguf-split : clarify operation of gguf-split (#19749 )	2026-03-25 13:12:50 +02:00
imatrix	common : move up common_init() and fix Windows UTF-8 logs (#21176 )	2026-03-31 12:53:41 +02:00
llama-bench	llama-bench: print `-n-cpu-moe` when offloaded layers > 1 (#20984 )	2026-03-25 21:17:27 +08:00
mtmd	model, mtmd: fix gguf conversion for audio/vision mmproj (#21309 )	2026-04-02 17:10:32 +02:00
parser	common/parser: fix call ID detection (Mistral parser mostly) + atomicity for tag-json parsers (#21230 )	2026-04-03 17:51:52 +02:00
perplexity	common : move up common_init() and fix Windows UTF-8 logs (#21176 )	2026-03-31 12:53:41 +02:00
quantize	llama : refactor llama_model_quantize_params to expose a pure C interface (#20346 )	2026-04-01 08:43:00 +03:00
results	common : move up common_init() and fix Windows UTF-8 logs (#21176 )	2026-03-31 12:53:41 +02:00
rpc	Fix locale-dependent float printing in GGUF metadata (#17331 )	2026-03-04 09:30:40 +01:00
server	server: save and clear idle slots on new task (`--clear-idle`) (#20993 )	2026-04-03 19:02:27 +02:00
tokenize	Fix locale-dependent float printing in GGUF metadata (#17331 )	2026-03-04 09:30:40 +01:00
tts	common : move up common_init() and fix Windows UTF-8 logs (#21176 )	2026-03-31 12:53:41 +02:00
CMakeLists.txt	llama: end-to-end tests (#19802 )	2026-03-08 12:30:21 +01:00