llama.cpp

History

itigges22 279e6c721e fix: CPU staging copy for recurrent state checkpoint (fixes crash) Root cause found: copy_cell crashes during find_slot because it calls ggml_backend_tensor_copy on GPU tensors while the compute graph is being built. Fixed by using CPU staging: tensor_get (GPU→CPU) then tensor_set (CPU→GPU). Also increased rs_size from 1 to 3 cells per sequence to make room for checkpoint cells needed by speculative decoding rollback. Results: - No more crashes during speculative decode - 23.8 tok/s with MTP (vs 16.7 without) - 75% acceptance rate - Output still garbled on long generation due to seq_rm not finding checkpoints at the right positions (checkpoint position mismatch) Next: fix checkpoint position tracking so seq_rm can find and restore the correct recurrent state after draft rejection.		2026-03-20 12:47:52 -04:00
..
batched-bench	Fix locale-dependent float printing in GGUF metadata (#17331 )	2026-03-04 09:30:40 +01:00
cli	tools/cli: fix disable reasoning (#20606 )	2026-03-15 22:40:53 +01:00
completion	chore : correct typos [no ci] (#20041 )	2026-03-05 08:50:21 +01:00
cvector-generator	chore : correct typos [no ci] (#20041 )	2026-03-05 08:50:21 +01:00
export-lora	Fix locale-dependent float printing in GGUF metadata (#17331 )	2026-03-04 09:30:40 +01:00
fit-params	llama-fit-params: keep explicit --ctx-size 0 (#19070 )	2026-01-24 22:13:08 +01:00
gguf-split	Fix locale-dependent float printing in GGUF metadata (#17331 )	2026-03-04 09:30:40 +01:00
imatrix	chore : correct typos [no ci] (#20041 )	2026-03-05 08:50:21 +01:00
llama-bench	llama-bench: introduce `-hf` and `-hff` flags & use `--mmap 1` by default (#20211 )	2026-03-09 09:05:44 +08:00
mtmd	mtmd: add llama-mtmd-debug binary (#20508 )	2026-03-14 15:52:29 +01:00
parser	Autoparser - complete refactoring of parser architecture (#18675 )	2026-03-06 21:01:00 +01:00
perplexity	tools : enable kvu in perplexity for hellaswag, winogrande, multiple-choice (#19954 )	2026-03-13 21:25:57 +01:00
quantize	llama-quant : fail early on missing imatrix, refactor type selection, code cleanup (#19770 )	2026-03-10 08:16:05 +02:00
results	llama: end-to-end tests (#19802 )	2026-03-08 12:30:21 +01:00
rpc	Fix locale-dependent float printing in GGUF metadata (#17331 )	2026-03-04 09:30:40 +01:00
server	fix: CPU staging copy for recurrent state checkpoint (fixes crash)	2026-03-20 12:47:52 -04:00
tokenize	Fix locale-dependent float printing in GGUF metadata (#17331 )	2026-03-04 09:30:40 +01:00
tts	Fix locale-dependent float printing in GGUF metadata (#17331 )	2026-03-04 09:30:40 +01:00
CMakeLists.txt	llama: end-to-end tests (#19802 )	2026-03-08 12:30:21 +01:00