llama.cpp/tools
itigges22 bc443d36a8 fix: MTP cooldown after draft rejection + debug logging
- Add cooldown flag to MTP speculative state: after draft rejection,
  skip next proposal to force single-token decode for fresh MTP logits
- Root cause: MTP logits are from the last batch position (draft token).
  When draft is rejected, next proposal uses stale/wrong logits (13% accept).
  With cooldown: proposals only use fresh single-token MTP logits (95% accept).
- Simplified seq_rm fallback: log and continue instead of re-evaluating
- Added debug logging (MTP-DBG, MTP-VERIFY) for acceptance rate tracking
- Results: 95% acceptance rate, 0 restarts, no garbled output on 2048 tokens
2026-03-19 15:30:01 -04:00
..
batched-bench Fix locale-dependent float printing in GGUF metadata (#17331) 2026-03-04 09:30:40 +01:00
cli tools/cli: fix disable reasoning (#20606) 2026-03-15 22:40:53 +01:00
completion chore : correct typos [no ci] (#20041) 2026-03-05 08:50:21 +01:00
cvector-generator chore : correct typos [no ci] (#20041) 2026-03-05 08:50:21 +01:00
export-lora Fix locale-dependent float printing in GGUF metadata (#17331) 2026-03-04 09:30:40 +01:00
fit-params llama-fit-params: keep explicit --ctx-size 0 (#19070) 2026-01-24 22:13:08 +01:00
gguf-split Fix locale-dependent float printing in GGUF metadata (#17331) 2026-03-04 09:30:40 +01:00
imatrix chore : correct typos [no ci] (#20041) 2026-03-05 08:50:21 +01:00
llama-bench llama-bench: introduce `-hf` and `-hff` flags & use `--mmap 1` by default (#20211) 2026-03-09 09:05:44 +08:00
mtmd mtmd: add llama-mtmd-debug binary (#20508) 2026-03-14 15:52:29 +01:00
parser Autoparser - complete refactoring of parser architecture (#18675) 2026-03-06 21:01:00 +01:00
perplexity tools : enable kvu in perplexity for hellaswag, winogrande, multiple-choice (#19954) 2026-03-13 21:25:57 +01:00
quantize llama-quant : fail early on missing imatrix, refactor type selection, code cleanup (#19770) 2026-03-10 08:16:05 +02:00
results llama: end-to-end tests (#19802) 2026-03-08 12:30:21 +01:00
rpc Fix locale-dependent float printing in GGUF metadata (#17331) 2026-03-04 09:30:40 +01:00
server fix: MTP cooldown after draft rejection + debug logging 2026-03-19 15:30:01 -04:00
tokenize Fix locale-dependent float printing in GGUF metadata (#17331) 2026-03-04 09:30:40 +01:00
tts Fix locale-dependent float printing in GGUF metadata (#17331) 2026-03-04 09:30:40 +01:00
CMakeLists.txt llama: end-to-end tests (#19802) 2026-03-08 12:30:21 +01:00