llama.cpp

Commit Graph

Author	SHA1	Message	Date
Jeff Bolz	2524c26164	vulkan: fix push constant size for quantize_q8_1 (#18687 ) I added an assert to catch further mismatches, and it found several. Fix those, too.	2026-01-08 15:40:58 +01:00
Jeff Bolz	cb14b06995	vulkan: optimize ssm_scan (#18630 ) * vulkan: optimize ssm_scan * fix warp vs subgroup naming	2026-01-08 15:16:54 +01:00
Adrien Gallouët	55abc39355	vendor : update cpp-httplib to 0.30.0 (#18660 ) * vendor : update cpp-httplib to 0.30.0 * common : allow custom headers when downloading	2026-01-08 13:53:54 +01:00
Georgi Gerganov	f2f6c88067	scripts : support chaining commands in pr2wt.sh (#18671 )	2026-01-08 13:40:23 +02:00
도로로도로또	945bf10627	metal : add MoE kernel specialization for ne20=5 (#18667 ) Add template specialization for kernel_mul_mm_id_map0 with ne20=5 to support models using 5 active experts (e.g., VAETKI).	2026-01-08 12:37:45 +02:00
Johannes Gäßler	64848deb18	llama-fit-params: free memory target per device (#18679 )	2026-01-08 10:07:58 +01:00
Doctor Shotgun	9a5724dee2	ggml: add env var GGML_OP_OFFLOAD_MIN_BATCH (#18535 ) * ggml: add env var GGML_OP_OFFLOAD_MIN_BATCH * makes the min_batch_size for triggering op offload configurable via env var, defaulting to the prior hardcoded value of 32 * ggml: read GGML_OP_OFFLOAD_MIN_BATCH once and store to dev ctx * cann: forward declaration of device context struct * cann: move offload op check after device context declaration * cuda: fix whitespace Co-authored-by: Aman Gupta <amangupta052@gmail.com> --------- Co-authored-by: Aman Gupta <amangupta052@gmail.com>	2026-01-08 11:03:21 +02:00
Daniel Bevenius	9c142e3a2a	model-conversion : add warn about transformers mismatch (#18691 ) This commit adds a check comparing the installed transformers library with the transformers version that the original model supports. This check will be performed upon a model verification failure and prints a warning/hint to the user suggesting to install the correct version of the transformers library. The motivation for this change is that it is possible for the model verification to fail due to differences in the transformers library used and it might not be obvious that this could be the cause of the failure. With this warning the correct version can be checked and hopefully save time troubleshooting the cause of the verification failure.	2026-01-08 09:29:53 +01:00
Daniel Bevenius	df7fb92170	model-conversion : remove -st targets for converted model (#18689 ) This commit removes the '-st` make target for running the converted embedding model. The motivation for this is that the pooling type is now part of the .gguf metdata of the model and this is used by llama-debug when running the model. So there is no need to specify the pooling type separately any more. The commit also adds an option to specify the type of normalization applied to the output embeddings when running the converted model. And the readme documentation has been updated to reflect these changes.	2026-01-08 09:29:15 +01:00
Julius Tischbein	2038101bd9	llama : add `use_direct_io` flag for model loading (#18166 ) * Adding --direct-io flag for model loading * Fixing read_raw() calls * Fixing Windows read_raw_at * Changing type off_t to size_t for windows and Renaming functions * disable direct io when mmap is explicitly enabled * Use read_raw_unsafe when upload_backend is available, not functional on some devices with Vulkan and SYCL * Fallback to std::fread in case O_DIRECT fails due to bad address * Windows: remove const keywords and unused functions * Update src/llama-mmap.cpp Co-authored-by: Georgi Gerganov <ggerganov@gmail.com> --------- Co-authored-by: jtischbein <jtischbein@gmail.com> Co-authored-by: Georgi Gerganov <ggerganov@gmail.com>	2026-01-08 08:35:30 +02:00
shaofeiqi	568371a726	opencl: add FILL op support (#18682 )	2026-01-07 22:04:50 -08:00
Sigbjørn Skjæret	5b8844ae53	scripts : fix repos cloned with .git extension (#18669 )	2026-01-07 22:35:34 +01:00
Sigbjørn Skjæret	7e16fef085	convert : more variants of rope_theta config entries (#18668 )	2026-01-07 22:34:51 +01:00
Oliver Walsh	f5245b5e4e	cuda : fix build on cuda 12.8 (#18672 ) compute121 requires 12.9 Signed-off-by: Oliver Walsh <owalsh@redhat.com>	2026-01-07 22:32:44 +01:00
R	ae9f8df778	fix(docker): add missing libglvnd libraries to Vulkan image (#18664 ) Add libglvnd0, libgl1, libglx0, libegl1, libgles2 to the Vulkan Dockerfile base image. These libraries are required by mesa-vulkan-drivers to properly initialize the Vulkan ICD and detect GPU devices. Without these libraries, vkEnumeratePhysicalDevices() returns an empty list, resulting in "ggml_vulkan: No devices found." error. Fixes #17761	2026-01-07 16:57:42 +01:00
Adrien Gallouët	56d2fed2b3	tools : remove llama-run (#18661 ) * tools : remove llama-run * Remove licenses/LICENSE-linenoise Signed-off-by: Adrien Gallouët <angt@huggingface.co>	2026-01-07 16:18:26 +01:00
Georgi Gerganov	56426673cb	scripts : add pr2wt.sh (#18644 ) * scripts : add pr2wt.sh * script : shebang Co-authored-by: Sigbjørn Skjæret <sigbjorn.skjaeret@scala.com> --------- Co-authored-by: Sigbjørn Skjæret <sigbjorn.skjaeret@scala.com>	2026-01-07 15:16:20 +02:00
Daniel Bevenius	bb77764c2d	convert : clarify sentence-transformers-dense-modules help [no ci] (#18662 ) * convert : clarify sentence-transformers-dense-modules help [no ci] This commit updates this options help message which currently looks like this: ```console --sentence-transformers-dense-modules Whether to include sentence-transformers dense modules.It can be used for sentence-transformers models, like google/embeddinggemma-300mDefault these modules are not included. ```	2026-01-07 13:18:53 +01:00
Sigbjørn Skjæret	9dfa8ee950	ci : run cann build unconditionally [no ci] (#18659 )	2026-01-07 13:07:08 +01:00
Jeff Bolz	ca4a8370bc	vulkan: reject ops when a tensor is too large to allocate (#18646 )	2026-01-07 12:03:32 +01:00
virajwad	03023296cf	vulkan: Warptile tuning for Intel Xe2/Xe3 (#18178 ) * modify warptile tuning for xe3 * intel vendor check w/ coopmat support * fix back formatting * fix formatting change 2 * move intel check to chip specific tuning part * Change to support both windows and linux * modify m_warptile to l_warptile for intel * modify warptile tuning for bf16 matmuls to fix regression (m_warptile to l_warptile) * Code style changes * Code style changes (2) * Code style changes (3)	2026-01-07 11:59:47 +01:00
Eve	8c77a04cc7	vulkan: more mul mat optimizations (#18533 ) * q4_k * q5_k * q2_k * q4_1 * q5_1 * better buf index	2026-01-07 11:13:17 +01:00
Daniel Bevenius	ffba4f29e6	examples : add debug utility/example (#18464 ) * examples : add debug utility/example This commit introduces a new example named llama-debug which is a utility that is intended to be used to assist with developing/debugging a converted model. The motivation for this utilitiy is to assist in model conversion work to verify that the model produces the expected outputs. It is intended to replace logits.cpp in examples/model-conversion. Example usage: ```console ./build/bin/llama-debug \ -m models/Qwen2.5-0.5B-Instruct.gguf \ --prompt "Hello, my name is" \ --save-logits ... Model add_bos: false Input prompt: "Hello, my name is" Token ids (5): Hello(9707) ,(11) my(847) name(829) is(374) Data saved to data/llamacpp-Qwen2.5-0.5B-Instruct.bin Data saved to data/llamacpp-Qwen2.5-0.5B-Instruct.txt Prompt saved to data/llamacpp-Qwen2.5-0.5B-Instruct-prompt.txt Tokens saved to data/llamacpp-Qwen2.5-0.5B-Instruct-tokens.bin ``` For more details about the options available for this example, please refer to examples/debug/README.md. * throw runtime error instead of logging error * remove params.warmup and enable the warmup/nowarmup option * model-conversion : remove logits.cpp This commit removes logits.cpp in favor of using llama-debug for generating logits and embeddings. * examples : remove model-conversion directory This was missed in the previous commit. * model-conversion : add support for saving prompt and token ids This commit add support for storing the prompt and the token ids for the prompt when running the original models. The motivation for this is that this will allow us to compare the prompt and the tokens generated for the prompt when verifing the converted model. Currently it is possible that even if the same prompt is used that the tokens generated are different if there is a difference in the tokenization between the original and converted model which would currently go unnoticed (the verification will most likely fail but it might not be obvious why). * squash! model-conversion : add support for saving prompt and token ids fix pyright errors. * model-conversion : add compare_tokens utility This commit adds a script to compare token outputs between original and converted models. Example usage: ```console (venv) $ ./scripts/utils/compare_tokens.py pytorch-gemma-3-270m-it llamacpp-gemma-3-270m-it-bf16 Comparing tokens between: Original : pytorch-gemma-3-270m-it (6 tokens) Converted: llamacpp-gemma-3-270m-it-bf16 (6 tokens) ✅ All 6 tokens match! ``` And there is a verbose flag that will also print out the prompts: ```console (venv) $ ./scripts/utils/compare_tokens.py pytorch-gemma-3-270m-it llamacpp-gemma-3-270m-it-bf16 -v Original model prompt (pytorch-gemma-3-270m-it): prompt: Hello, my name is n_tokens: 6 token ids: 2, 9259, 236764, 1041, 1463, 563 Converted model prompt (llamacpp-gemma-3-270m-it-bf16): prompt: Hello, my name is n_tokens: 6 token ids: 2, 9259, 236764, 1041, 1463, 563 Comparing tokens between: Original : pytorch-gemma-3-270m-it (6 tokens) Converted: llamacpp-gemma-3-270m-it-bf16 (6 tokens) ✅ All 6 tokens match! ``` * model-conversion : add token comparison to verifiction scripts This commit add the calling of the compare_tokens function in compare-logits.py and semantic_check.py to ensure that the token ids that the tokenizers procoduce are the same before proceeding with verifying the logits/embeddings. Placing them in the existing scripts instead calling them separately ensures that the token comparison is always done prior to the logit/embedding verifications. Follow up commit/pr could refactor the causal logits verification into a single script instead of the two that exist now. This would reduce the code and make it consistent with the embeddings verficiation which only has a single script. * debug : use llama_model_n_embd_out This commit updates the debug example to use the new function llama_model_n_embd_out instead of llama_model_n_embd. The motivation for this change is to support late interation retriever models, like LFM2-ColBert-350M, where the output embeddings are down projected to a lower dimension. * debug : add print_usage function This commit adds a print_usage function that is passed to the common_params_parse. The motivation for this is that this enables a specific usage message which will be printed after all the options, for example: ```console example usage: Print tensors: ./build/bin/llama-debug -m model.gguf -p "Hello my name is" --verbose The tensors to be printed can be filtered with --tensor-filter option. Save logits/embeddings: ./build/bin/llama-debug -m model.gguf -p "Hello my name is" --save-logits Add --embedding to save embeddings ```	2026-01-07 10:42:19 +01:00
hipudding	3333951d86	CANN: Fix rename for get_env (#18652 ) In #18624, get_env in ggml-cann was renamed to get_env_as_lowercase to accurately reflect the function’s behavior and reduce the chance of misuse. However, the update missed renaming call sites in other files. This commit fixes that oversight.	2026-01-07 16:11:31 +08:00
Raul Torres	193ee38a1b	CANN: Rename `get_env` to `get_env_as_lowercase` (#18624 )	2026-01-07 10:01:25 +08:00
Max Krasnyansky	95ea9e0861	Hexagon add support for f16/f32 flash attention, scale, set-rows and improve f16/32 matmul (#18611 ) * hexagon: improve fp16 matmul and add fp32/fp16 flash-attention * hexagon: add support for set-rows fp32 -> fp16 with i32/i64 row-idx * hexagon: add support for SCALE fp32 * hexagon: replace scalar fp32 -> fp16 copy with HVX * hexagon: optimize flash_atten_ext with aligned VTCM buffers and DMA - Implements double-buffered DMA prefetching for K, V, and Mask tensors. - Ensures K and V rows in VTCM are padded to 128 bytes to support aligned HVX operations. - Correctly synchronizes DMA transfers to prevent race conditions. - Uses `FLASH_ATTN_BLOCK_SIZE` of 128 for efficient chunking. * hexagon: use aligned mad_f16 * hexagon: flash_atten more aligned ops * hexagon: optimize scale_f32 hvx helpers * hexagon: unroll fa loops * hexagon: remove unused set-rows log * hexagon: flash_attn_ext add support for DMAing Q - Update `op_flash_attn_ext` to include Q row size in scratchpad allocation. - Pad Q row size to 128 bytes for alignment. - Implement DMA transfer for Q tensor in `flash_attn_ext_f16_thread`. - Update dot product computations to use VTCM-buffered Q data. * hexagon: fix handling of NANs hvx dotproducts * hexagon: cleanup spad allocation in flash-atten * hexagon: improve fp16/fp32 matmul - Introduced `vec_dot_f16_f16` and `vec_dot_f16_f16_rx2` kernels using efficient HVX dot product intrinsics. - Added `quantize_fp32_f16` to copy/convert weights from DDR to VTCM - Updated `op_matmul` to use the optimized path when VTCM capacity allows and broadcasting requirements are compatible. - Implemented fallback logic to the original implementation for complex broadcasting scenarios. * hexagon: fix HVX_ARCH check * hexagon: matmul cleanup and fp16 fixes Use aligned vec_dot_f16 for 2d matmuls and unaligned version for 4d. * hexagon: fix fp16 x fp16 matmuls and some minor refactoring * hexagon: add support for GET_ROWS f32 -> f32 Also optimize SET_ROWS threading a bit when we have just a few rows to process. * hexagon: optimize set-rows threading * hexagon: update adb/run-bench.sh to properly support experimental and verbose options * hexagon: flash_atten use aligned vectors for dot products	2026-01-06 17:38:29 -08:00
Tarek Dakhran	ccbc84a537	mtmd: mtmd_audio_streaming_istft (#18645 ) Change is decoupled from https://github.com/ggml-org/llama.cpp/pull/18641. [LFM2.5-Audio-1.5B](https://huggingface.co/LiquidAI/LFM2.5-Audio-1.5B) needs streaming istft for generating output audio. * add streaming ISTFT class (`mtmd_audio_streaming_istft`) with overlap-add for audio reconstruction * replace global audio cache with per-instance cache, the model requires two independent caches, for preprocessing (audio input) and for istft (audio output). * unified templated FFT/IFFT implementation supporting both forward and inverse transforms	2026-01-06 21:00:29 +01:00
Johannes Gäßler	68b4d516c3	llama-params-fit: fix last devices with low VRAM (#18494 )	2026-01-06 20:02:30 +01:00
Aadeshveer Singh	24af22fc36	ggml : optimize cuda ssm_scan using warp-level reduction (#18505 ) * ggml : optimize cuda ssm_scan using warp-level reduction * ggml : apply code review suggestions (style, const, constexpr) * ggml : add TODO regarding stride consistency	2026-01-07 02:24:34 +08:00
Xuan-Son Nguyen	07fbe19f1f	arg: use CSV escape style for multiple-value args (#18643 ) * arg: use CSV escape style for multiple-value args * add test	2026-01-06 17:51:08 +01:00
Jeff Bolz	ea13cba850	vulkan: support buffer_from_host_ptr (#18467 ) * vulkan: support buffer_from_host_ptr * hacky use of buffer_from_host_ptr for directio * disable buffer_from_host_ptr cap * use external memory for ggml_vk_host_malloc, revert model loader changes * disable external_memory_host for MoltenVK * take buffer memory types into account * don't use external_memory_host for ggml_vk_host_malloc	2026-01-06 17:37:07 +01:00
Aman Gupta	090b137e56	ggml-cuda: refactor cuda graph usage (#18637 ) * ggml-cuda: refactor cuda graph usage * use is_enabled() instead of enabled	2026-01-06 23:48:45 +08:00
Beinsezii	968929528c	mmq.cu: tune mmq/rocblas switching for RDNA (#18537 ) * Patch perf regression for mmq kernels in ROCm recover performance regression for https://github.com/ggml-org/llama.cpp/issues/17917 * add n_experts branch like the cdna path * mmq.cu: tune mmq/wmma switching for RDNA * mmq.cu: move amd wmma mmq/wmma switching behind IS_RDNA3 * Update ggml/src/ggml-cuda/mmq.cu Co-authored-by: Johannes Gäßler <johannesg@5d6.de> --------- Co-authored-by: Jiacheng (Jason) Chen <76919340+jiachengjason@users.noreply.github.com> Co-authored-by: jiachengjason <jasonchen.jiacheng@gmail.com> Co-authored-by: Johannes Gäßler <johannesg@5d6.de>	2026-01-06 16:26:07 +01:00
R	3d26a09dc7	server : add thinking content blocks to Anthropic Messages API (#18551 ) * server : add thinking content blocks to Anthropic Messages API Add support for returning reasoning/thinking content in Anthropic API responses when using models with --reasoning-format deepseek and the thinking parameter enabled. - Non-streaming: adds thinking block before text in content array - Streaming: emits thinking_delta events with correct block indices - Partial streaming: tracks reasoning state across chunks via anthropic_has_reasoning member variable Tested with bartowski/DeepSeek-R1-Distill-Qwen-7B-GGUF model. * server : fix Anthropic API streaming for thinking content blocks Add signature field and fix duplicate content_block_start events in Anthropic Messages API streaming responses for reasoning models. * server: refactor Anthropic streaming state to avoid raw pointer Replace raw pointer to task_result_state with direct field copies: - Copy state fields in update() before processing chunk - Use local copies in to_json_anthropic() instead of dereferencing - Pre-compute state updates for next chunk in update() This makes the data flow clearer and avoids unsafe pointer patterns.	2026-01-06 16:17:13 +01:00
Christian Kastner	bd2a93d475	gguf-py : add requests to dependencies (#18629 )	2026-01-06 08:56:38 +01:00
Adrien Gallouët	e75ee11024	ggml : fix avx512bf16 build (#18623 ) - include `immintrin.h` when required - remove unused m512bh Signed-off-by: Adrien Gallouët <angt@huggingface.co>	2026-01-06 08:54:10 +02:00
Raul Torres	da9b8d3300	CANN: Make `valid_values` variable `static const` (#18627 )	2026-01-06 11:53:28 +08:00
nwyin	e443fbcfa5	ggml webgpu: add CEIL operation support (#18605 ) * ggml-webgpu: add CEIL operation support Add support for the CEIL unary operation in the WebGPU backend: - Add CEIL_FUNC shader template in unary_op.wgsl - Add 4 shader variants (f32, f16, inplace versions) - Initialize CEIL pipelines in ggml-webgpu.cpp - Register CEIL in supports_op function * docs: update WebGPU ops support for CEIL	2026-01-05 11:38:57 -08:00
Tarek Dakhran	73d284a250	model : add LFM2-ColBert-350M (#18607 ) * model : add LFM2-ColBert-350M * llama_model_n_embd_out() - returns `hparams.n_embd_out` if set and fallbacks to `hparams.n_embd`	2026-01-05 19:52:56 +01:00
Johannes Gäßler	df17a4c94f	CUDA: fix FA FP16 accumulator overflow for Granite (#18614 )	2026-01-05 19:51:13 +01:00
tt	1871f0ba56	add YoutuVLForConditionalGeneration architectures (#18620 ) * Support Youtu-VL Model --------- Co-authored-by: Xuan-Son Nguyen <son@huggingface.co> Co-authored-by: Sigbjørn Skjæret <sigbjorn.skjaeret@scala.com>	2026-01-05 18:15:14 +01:00
Aman Gupta	f47edb8c19	ggml-cuda: check for srcs outside the cgraph (#18583 ) * ggml-cuda: check for srcs outside the cgraph * review: use leafs instead	2026-01-05 22:46:36 +08:00
Vladislav Sayapin	da143b9940	server : fix router child env in containerized environments (#18562 )	2026-01-05 14:12:05 +01:00
Jeff Bolz	f1768d8f03	vulkan: fix topk_moe_sigmoid_norm_bias failures in GLM-4.6 (#18582 )	2026-01-05 11:51:39 +01:00
Georgi Gerganov	2da64a2f8a	models : fix backend assignment for Granite/Nemotron graphs (#18599 ) * models : fix backend assignment for Granite/Nemotron graphs * cont : add ref * cont : move call to build_inp_embd()	2026-01-05 12:34:23 +02:00
Jeff Bolz	b37124d2d2	vulkan: handle quantize_q8_1 overflowing the max workgroup count (#18515 ) * vulkan: handle quantize_q8_1 overflowing the max workgroup count * vulkan: Fix small tile size matmul on lavapipe * fix mul_mat_id failures	2026-01-05 11:30:14 +01:00
Sigbjørn Skjæret	eadc4184ca	llama : refactor rope_freq_base/scale_swa conversion and init (#18553 ) * refactor rope_freq_base/scale_swa conversion and init * safe defaults for unknowns * update relevant models * grammar * add get_rope_freq_scale to modern-bert * const * const * log swa info	2026-01-05 09:14:04 +01:00
Chenguang Li	67e3f6f601	CANN: add operator fusion support for ADD + RMS_NORM (#17512 ) This commit implements operator fusion for ADD + RMS_NORM operations in the CANN backend to reduce memory access overhead and improve performance. The fusion is controlled by the GGML_CANN_OPERATOR_FUSION environment variable (default: false). Changes: - Implement ggml_cann_op_add_rms_norm_fused() using ACLNN AddRmsNorm - Add ggml_cann_can_fuse() to check fusion eligibility - Integrate fusion logic into computation graph evaluation - Add test cases for ADD + RMS_NORM fusion - Update documentation with new environment variable The fusion combines ADD and RMS_NORM into a single kernel call, which is more efficient than executing them separately.	2026-01-05 15:38:18 +08:00
Francisco Herrera	92ac1e016b	doc: clarify that steps also apply to linux for opencl (#18002 ) * Clarify setup steps for Linux Added note that setup steps apply to Linux as well. * Added note for backtick replacement * clarify that backtick replacement only applies on linux * clarified Linux specific steps So actually some changes are needed for Linux but they are minor. * clarify change execution * clarify by placing info after steps * clarify which steps * Make instructions consistent across OSes * Rm whitespace * Update docs/backend/OPENCL.md Co-authored-by: Aaron Teo <taronaeo@gmail.com> * Update docs/backend/OPENCL.md Co-authored-by: Aaron Teo <taronaeo@gmail.com> * Update docs/backend/OPENCL.md Co-authored-by: Aaron Teo <taronaeo@gmail.com> --------- Co-authored-by: Aaron Teo <taronaeo@gmail.com>	2026-01-04 20:39:25 -08:00
Ali Tariq	8e3a761189	ci : init git lfs in every build for RISC-V (#18590 ) * Initialized git lfs in every test * Added git-lfs in dependencies to instal	2026-01-05 02:18:33 +01:00

1 2 3 4 5 ...

7677 Commits All Branches Search

7677 Commits

All Branches