llama.cpp

History

itigges22 19fdba56b5 feat: MTP support for dense Qwen 3.5 with FastMTP vocabulary trimming Add Multi-Token Prediction (MTP) speculative decoding for Qwen3.5 dense models (0.8B-27B). The MTP head uses a full transformer block (attention + FFN) to predict the next-next token, enabling ~28 tok/s on RTX 5060 Ti. Key changes: - Model loading: Qwen3.5 MTP layer tensors (nextn.eh_proj, attention weights, FFN) loaded into layers[n_layer-1] - Graph builder: Full MTP head with self-attention, gated RoPE, FFN, and vocabulary projection. Unfiltered hidden state passed for proper KV cache population during prompt processing. - FastMTP: Vocabulary trimming from 248K to 32K tokens via ggml_view_2d on the lm_head. Reduces draft generation from 22ms to 6ms (3.7x). - Speculative framework: MTP auto-detection for hybrid models, fuzzy seq_rm checkpoint matching for DeltaNet rollback. - Server: Two-phase decode option for hybrid/recurrent models to avoid DeltaNet state corruption from rejected drafts. - Recurrent state: Fixed copy_cell (ggml_view_1d takes element count, not bytes), buffer assignment for no_alloc views. Results on Qwen3.5-9B Q4_K_M (RTX 5060 Ti 16GB): - 28.1 tok/s with 82% acceptance rate (temp=0) - 92% acceptance with two-phase decode (correct output, 15 tok/s) - Draft generation: 6.1ms with FastMTP (vs 22.4ms full vocab)		2026-03-21 14:18:40 -04:00
..
batched-bench	Fix locale-dependent float printing in GGUF metadata (#17331 )	2026-03-04 09:30:40 +01:00
cli	common/parser: add proper reasoning tag prefill reading (#20424 )	2026-03-19 16:58:21 +01:00
completion	common/parser: add `--skip-chat-parsing` to force a pure content parser. (#20289 )	2026-03-17 16:16:43 +01:00
cvector-generator	chore : correct typos [no ci] (#20041 )	2026-03-05 08:50:21 +01:00
export-lora	Fix locale-dependent float printing in GGUF metadata (#17331 )	2026-03-04 09:30:40 +01:00
fit-params	llama-fit-params: keep explicit --ctx-size 0 (#19070 )	2026-01-24 22:13:08 +01:00
gguf-split	Fix locale-dependent float printing in GGUF metadata (#17331 )	2026-03-04 09:30:40 +01:00
imatrix	chore : correct typos [no ci] (#20041 )	2026-03-05 08:50:21 +01:00
llama-bench	llama-bench: introduce `-hf` and `-hff` flags & use `--mmap 1` by default (#20211 )	2026-03-09 09:05:44 +08:00
mtmd	mtmd: add clip_graph::build_mm() (#20751 )	2026-03-19 13:11:39 +01:00
parser	common/parser: add proper reasoning tag prefill reading (#20424 )	2026-03-19 16:58:21 +01:00
perplexity	tools : enable kvu in perplexity for hellaswag, winogrande, multiple-choice (#19954 )	2026-03-13 21:25:57 +01:00
quantize	llama-quant : fail early on missing imatrix, refactor type selection, code cleanup (#19770 )	2026-03-10 08:16:05 +02:00
results	llama: end-to-end tests (#19802 )	2026-03-08 12:30:21 +01:00
rpc	Fix locale-dependent float printing in GGUF metadata (#17331 )	2026-03-04 09:30:40 +01:00
server	feat: MTP support for dense Qwen 3.5 with FastMTP vocabulary trimming	2026-03-21 14:18:40 -04:00
tokenize	Fix locale-dependent float printing in GGUF metadata (#17331 )	2026-03-04 09:30:40 +01:00
tts	Fix locale-dependent float printing in GGUF metadata (#17331 )	2026-03-04 09:30:40 +01:00
CMakeLists.txt	llama: end-to-end tests (#19802 )	2026-03-08 12:30:21 +01:00