llama.cpp

Commit Graph

Author	SHA1	Message	Date
Jan Boon	68505e0686	Merge `23b5a5c026` into `e06088da0f`	2026-02-08 21:20:52 +08:00
Oliver Simons	e06088da0f	CUDA: Fix non-contig rope (#19338 ) * Rename variables + fix rope_neox Seems memory layout is shared with Vulkan so we can port fix from https://github.com/ggml-org/llama.cpp/pull/19299 * Fix rope_multi * Fix rope_vision * Fix rope_norm * Rename ne* to ne0* for consistent variable naming * cont : consistent stride names --------- Co-authored-by: Georgi Gerganov <ggerganov@gmail.com>	2026-02-08 15:12:51 +02:00
Adrien Gallouët	5fa1c190d9	rpc : update from common.cpp (#19400 ) Signed-off-by: Adrien Gallouët <angt@huggingface.co>	2026-02-08 09:06:45 +01:00
Georgi Gerganov	eb449cdfa4	server : improve context checkpoint logic (#19408 )	2026-02-08 09:40:04 +02:00
ddh0	5999b50eb0	llama-quantize : cleanup `--help` output (#19317 ) * cleanup `llama-quantize --help` output some much needed TLC * remove future argument oops, spoiler * cleanup of cleanup	2026-02-08 09:22:38 +02:00
Sigbjørn Skjæret	9a5f57795c	ci : remove server job from webui and move slow test (#19424 ) * remove server job from webui and move slow test * use pip-install option	2026-02-08 01:20:00 +01:00
Georgi Gerganov	96441c955e	ci : use -j param correctly when building with sanitizers (#19411 ) * ci : use less jobs when building with sanitizers * cont : fix nproc * cont : fix the fix * cont : simplify	2026-02-07 23:50:47 +01:00
Jan Boon	23b5a5c026	sampling : fix whitespace	2026-02-07 09:19:23 +00:00
Georgi Gerganov	8872ad2125	metal : consolidate bin kernels (#19390 ) * metal : refactor bin kernels * cont * cont : fix cv	2026-02-07 10:35:56 +02:00
Georgi Gerganov	34ba7b5a2f	metal : fix event synchronization in cpy_tensor_async (#19402 )	2026-02-07 07:37:15 +02:00
Jan Boon	a0323a989d	sampling : comment on state use	2026-02-07 05:03:21 +00:00
Jan Boon	15ade86a75	sampling : simplify clone	2026-02-07 04:49:37 +00:00
Jan Boon	1b1b2cbe0e	sampling : also apply sorting in backend path when blue noise rng is selected	2026-02-07 04:27:52 +00:00
Jan Boon	267cd808a2	sampling : cleanup blue noise rng with some more notes	2026-02-07 02:42:21 +00:00
Jan Boon	6f7f63ee4d	merge	2026-02-07 10:09:23 +08:00
forforever73	b83111815e	model : support Step3.5-Flash (#19283 ) * Support Step3.5-Flash * fix: norm.weight + 1 (HF zero_centered=true) * step35: simplify GGUF conversion + drop redundant rope KVs * Address review feedback * rename limits -> clamp * Apply suggestions from code review Co-authored-by: Sigbjørn Skjæret <sigbjorn.skjaeret@scala.com> * Apply suggestion from @CISC Co-authored-by: Sigbjørn Skjæret <sigbjorn.skjaeret@scala.com> * rename swiglu limits -> swiglu clamp in LLM_KV * avoid CI fail * Apply suggestions from code review * Apply suggestions from code review * disabled KV shifting for LLM_ARCH_STEP35 * Apply suggestions from code review * mistakenly removed cmath * add model size && apply missed suggestion * assert partial_rotary_factors * fix CI errors: * load freq_base_swa --------- Co-authored-by: lvyichen <lvyichen@stepfun.com> Co-authored-by: Sigbjørn Skjæret <sigbjorn.skjaeret@scala.com>	2026-02-06 21:06:14 +01:00
Alex Trotta	3228e77287	gguf-py : bump sentencepiece version (#19319 ) * gguf-py: Bump sentencepiece version There's a new version that's been out for a while that addresses the issues mentioned in https://github.com/ggml-org/llama.cpp/pull/14200. There's a long chain of reasons I would like this change, but the short version is that it allows people who use both `sentencepiece` and `gguf` to take advantage of these fixes. On conda-forge, currently, it locks the version (since there is no notion of optional dependencies). Regardless, I don't think this should be too controversial. * review feedback	2026-02-06 21:05:19 +01:00
Abhijit Ramesh	7fbd36c50c	ggml-webgpu: JIT compile binary operators and handle binding overlaps (#19310 ) * ggml webgpu: port binary operators to use pre-wgsl * Add binary.wgsl: unified shader with conditionals for all 4 ops * Add gen_binary_shaders.cpp: build tool for using pre_wgsl preprocessor * Remove bin_op.tmpl.wgsl and binary.wgsl (Python template) * Update CMake to generate binary operator shaders at build time * ggml-webgpu: migrate binary ops to JIT compilation with overlap handling * port binary operators from AOT to pre-wgsl JIT compilation * add src1=dst overlap handling for binary ops * use compile-time workgroup size defines instead of runtime overrides * ggml-webgpu: complete overlap handling for binary ops * add support for inplace & overlap case in binding setup * restructure conditional logic to handle all overlap cases * ensure all buffer bindings are correctly assigned for edge cases * ggml-webgpu: remove unused binary overlap cases Remove src0==src1 binary overlap case that never occurs in practice. * keep INPLACE (src0==dst), OVERLAP (src1==dst), DEFAULT * remove unused src0==src1 and all-same variant * refactor wgsl to eliminate duplication	2026-02-06 10:33:30 -08:00
Nechama Krashinski	537eadb1b9	sycl: add F16 support for GGML_OP_CEIL (#19306 ) * Fix SYCL CEIL operator * sycl: implement GGML_OP_CEIL	2026-02-06 23:13:44 +08:00
Jeff Bolz	db6adb3c88	tests: reduce number of FA test permutations (#19381 ) Only test non-F16 for head size 64 and 72 (one a multiple of QK, one not).	2026-02-06 08:50:30 -06:00
Georgi Gerganov	dfde5993ea	common : add common_speculative_is_compat() (#19270 ) * llama : add llama_memory_can_rm_suffix() * Revert "llama : add llama_memory_can_rm_suffix()" This reverts commit `d30e59b62a`. * spec : check if the target context is compatible for spec decoding	2026-02-06 16:47:22 +02:00
Lasse Lauwerys	06bf3796f4	unicode : MSVC regex fix (#19340 ) * Fix model loading regex error * Change comments * Use const_iterator and remove specializations --------- Co-authored-by: Alde Rojas <hello@alde.dev>	2026-02-06 15:56:13 +02:00
ymcki	3688c4f504	Kimi-Linear support (backend agnostic + MLA KV cache) (#18755 ) * kimi linear model implementation * kimi linear convert_hf_to_gguf * kimi linear constants.py tensor_mapping.py * Kimi Linear ggml.h * kimi linear ggml-cpu * Kimi Linear ggml-cuda * Kimi Linear ggml.c * kimi linear src/llama * remove "const int64_t n_seq_tokens = q->ne[2];" to get rid of unused variable warning * remove type mismatch warning * read MoE params * removed some hard coded code * removed all hard code * use DeepseekV2 tokenizer * removed unnecessary internal methods called by the old set_vocab of KimiLinear * rewrite get_vocab for KimiLinear. Removed all kda_scan code * removed all traces of kda_scan * reduce OP count by 1 due to removal of kda_scan * Move KIMI_LINEAR to llm_arch_is_hybrid to enable KV cache * set n_embd_head_k/v to ensure kv cache works * don't quantize conv1d of Kimi Linear * Kimi Linear backend agnostic * removed LOG_INFO * naive chunking form implemented * fixed some comments * add Kimi-K2 specific tokens to be recognized as EOG * build_kda_autoregressive is implemented to replace build_kda_recurrent for faster inference. sync'd to b7682 * replaced Akk and Aqk with mul_mat and clamp * no clamp version * Moved Aqk computation out of the loop * fixed typo and split wkv_b into wk_b and wv_b * MLA KV cache support * fix trailing spaces * moved const llama_model & model; around to follow qwen3next format and see if it cna pass the -Wunused-private-field error * fix trailing whitespace * removed traling whitespaces in empty line + make sure indentation is multiple of 4 * try to make lint happy * remove blank lines to make lint happy * removed at least blank line containing white space * fixed flake8 complaints locally * return ggml_tensor * pair in kda_autoregressive and kda_chunking as in ngxson's Qwen3Next improvement * removed Kimi-Linear specific change that causes failure at server-windows * removed private: from kimi_linear to make build checks happy * removed unnecessary ggml_cont before ggml_reshape * created static function causal_conv1d to abtract similar code for q/k/v * merged dt_bias to SSM_DT. Do -exp(log_A) in convert_hf_to_gguf.py. * reverted to original * fixed find_hparam calls. Fixed e_score_correction_bias to use bias instead of weight. Removed all ssm_conv bias terms. * remove DT_B from constants.py. remove one comment line in llama-model.cpp * new class llm_graph_input_mem_hybrid_k to get around the new MLA change. switch the concat order of ggml_concat calls in kimi-linear.cpp to accommodate MLA changes. Removed support for exp_probs_b.weight * remove ssm_o_norm_b * remove ssm_o_norm_b * changed hparams.kda_head_dim to hparams.n_embd_head_kda. added TODO comment for class llama_graph_mem_hybrid_k * removed all ggml_cont b4 ggml_reshape_4d * Whitespace * replaced all hparams.get with find_hparams * added new names for n_experts, n_experts_used and score_func in TextModel and removed their code in KimiLinear in convert_hf_to_gguf.py. Removed unnecessary ggml_cont and GGML_ASSERT in kimi-linear.cpp * use is_mla to switch between different mem_hybrid types * fixed logical errors in convert_hf_to_gguf.py pointed out by CISC * removed if else for required parameters kv_lora_rank and qk_rope_head_dim * add back ggml_cont for Vcur * minor changes * removed extra line in llama-vocab.cpp. Added back the comment in llama-graph.cpp * f16 gguf cannot run without context length * made a mistake of adding back n_ctx parsing --------- Co-authored-by: Piotr Wilkin (ilintar) <piotr.wilkin@syndatis.com>	2026-02-06 11:39:58 +01:00
Jeff Bolz	1946e46f4c	vulkan: For coopmat2 FA, use fp16 accumulators for the final result (#19376 ) The cpu and cuda backends use fp16 for the VKQ accumulator type, this change does the same for vulkan. This helps particularly with large head sizes which are very register-limited. I tried this for the coopmat1 path and it slowed down a bit. I didn't try for scalar. I applied the softmax bias that the cuda backend uses to avoid overflow, although I was not able to reproduce the original bug without it.	2026-02-06 09:15:13 +01:00
Jeff Bolz	f9bd518a6b	vulkan: make FA mask/softcap enables spec constants (#19309 ) * vulkan: make FA mask/softcap enables spec constants * don't specialize for sinks * bump timeout a little bit	2026-02-06 08:49:58 +01:00
Georgi Gerganov	7fcf1ef45d	metal : skip loading all-zero mask (#19337 ) * metal : skip loading all-zero mask * cont : minor	2026-02-06 09:25:11 +02:00
Daniel Bevenius	e696cfc016	llama : rename llama-sampling to llama-sampler (#19363 ) This commit addresses the TODO in llama-sampling.h to rename that header and the implementation to llama-sampler.	2026-02-06 07:26:54 +01:00
Georgi Gerganov	3e21647666	cuda : cuda graphs now compare all node params (#19383 )	2026-02-06 07:55:06 +02:00
Georgi Gerganov	22cae83218	metal : adaptive CPU/GPU interleave based on number of nodes (#19369 )	2026-02-05 19:07:22 +02:00
Jeff Bolz	449ec2ab07	vulkan: Preprocess FA mask to detect all-neg-inf and all-zero. (#19281 ) Write out a 2-bit code per block and avoid loading the mask when it matches these two common cases. Apply this optimization when the mask is relatively large (i.e. prompt processing).	2026-02-05 09:26:38 -06:00
Georgi Gerganov	3795cc1e89	benches : update models + numbers (#19359 ) * bench : update script * benches : update numbers	2026-02-05 14:34:07 +02:00
Jan Boon	e829f2904e	sampling : blue noise requires tokens to be sorted	2026-02-05 10:46:18 +00:00
Sigbjørn Skjæret	b828e18c75	docker : fix vulkan build (#19352 )	2026-02-05 11:10:39 +01:00
Adrien Gallouët	a4ea7a188f	vendor : update BoringSSL to 0.20260204.0 (#19333 ) Signed-off-by: Adrien Gallouët <angt@huggingface.co>	2026-02-05 09:53:35 +01:00
Georgi Gerganov	7a4f97d196	metal : add diag (#19330 )	2026-02-05 10:08:45 +02:00
Oleksandr Kuvshynov	a498c75ad1	vulkan: fix GPU deduplication logic. (#19222 ) * vulkan: fix GPU deduplication logic. As reported in https://github.com/ggml-org/llama.cpp/issues/19221, the (same uuid, same driver) logic is problematic for windows+intel igpu. Let's just avoid filtering for MoltenVK which is apple-specific, and keep the logic the same as before `88d23ad5` - just dedup based on UUID. Verified that MacOS + 4xVega still reports 4 GPUs with this version. * vulkan: only skip dedup when both drivers are moltenVk	2026-02-05 09:06:59 +01:00
Jeff Bolz	3409ab842d	vulkan: Set k_load_shmem to false when K is too large (#19301 )	2026-02-05 08:48:33 +01:00
Jeff Bolz	c342c3b93d	vulkan: fix non-contig rope (#19299 )	2026-02-05 08:38:59 +01:00
will-lms	af252d0758	metal : add missing includes (#19348 )	2026-02-05 08:05:09 +02:00
Jan Boon	d5def78bb0	llama : note on blue noise	2026-02-05 03:15:55 +00:00
Sigbjørn Skjæret	11fb327bf3	vendor : add missing llama_add_compile_flags (#19322 ) * add missing llama_add_compile_flags * disable all warnings for ssl, crypto and fipsmodule	2026-02-05 02:27:38 +01:00
Jan Boon	ad73188337	llama : note on blue noise properties	2026-02-05 00:53:51 +00:00
Jan Boon	766d86df29	llama : cleanup and restore alternate code path	2026-02-04 23:35:22 +00:00
Jan Boon	3b4061981b	llama : make the sampler rng modular	2026-02-04 23:05:54 +00:00
Aaron Teo	e6e934c5ea	vendor: update cpp-httplib version (#19313 ) Signed-off-by: Aaron Teo <aaron.teo1@ibm.com>	2026-02-05 05:15:03 +08:00
Daniel Bevenius	b536eb0233	codeowners : add danbev for examples/debug (#19332 ) * codeowners : add danbev for examples/debug * Add @pwilkin to CODEOWNERS for debug --------- Co-authored-by: Piotr Wilkin (ilintar) <piotr.wilkin@syndatis.com>	2026-02-04 20:20:40 +01:00
Jan Boon	f271576d81	llama : initial blue noise test implementation	2026-02-04 19:12:50 +00:00
Jan Boon	b2ee2fbc0a	llama : add floating point blue noise rng	2026-02-04 18:18:56 +00:00
Jan Boon	e856c8f959	llama : add blue noise rng implementation	2026-02-04 17:46:44 +00:00
Xuan-Son Nguyen	e0c93af2a0	debug: make common_debug_print_tensor readable (#19331 ) * debug: make common_debug_print_tensor readable * editorconfig	2026-02-04 17:55:31 +01:00

1 2 3 4 5 ...

7987 Commits All Branches Search

7987 Commits

All Branches