llama.cpp

Commit Graph

Author	SHA1	Message	Date
ryan-mangeno	6d86944cb4	working through previous attemp, implimented more accurate conversion per previous attempt, added local sliding window attention that alternates every third layer	2025-09-03 14:32:39 -04:00
ryan-mangeno	ca353d37b4	fixed pre tokenizer and still working through previous pr	2025-09-02 12:26:20 -04:00
ryan-mangeno	c73eb685fd	added cls token per previous modern bert attempt, still working on checking out the rest	2025-08-29 12:15:31 -04:00
ryan-mangeno	2a1c75047c	ubatch issues, the assert for checking equal seqs in llama-graph.cpp when building attention keeps failing, setting ubatch size to 1 when running llama-embedding with --ubatch-size 1 makes it work, but needs to be looked into more	2025-08-28 12:59:42 -04:00
ryan-mangeno	853f344cfe	more cleanup	2025-08-28 12:47:10 -04:00
ryan-mangeno	40249dd5ec	cleanup	2025-08-28 12:37:02 -04:00
ryan-mangeno	9805635c12	cleanup	2025-08-28 12:36:26 -04:00
ryan-mangeno	8f328431a1	cleanup	2025-08-28 12:33:52 -04:00
ryan-mangeno	bffe3c9092	tensor debugging now works -> (llama-eval-callback), instead of simulated gate split with views, GEGLU is now used which does exactly this	2025-08-28 11:15:10 -04:00
ryan-mangeno	18c0c23ed8	fixed tensor mappings and working on buildin graph	2025-08-27 15:32:20 -04:00
ryan-mangeno	4ceb828112	correct tensor shape for qkv	2025-08-26 13:03:14 -04:00
ryan-mangeno	cc3d7abab4	continuing	2025-08-26 12:38:38 -04:00
ryan-mangeno	41b6864333	cleanup	2025-08-26 12:33:11 -04:00
ryan-mangeno	cc40378d27	some cleanup	2025-08-25 16:31:08 -04:00
ryan-mangeno	ac67fc6887	working on support, now working on building graph	2025-08-25 16:15:40 -04:00
ryan-mangeno	6643c5a852	conversion now working, hf -> gguf	2025-08-21 12:42:32 -04:00
ryan-mangeno	6151592ea7	constants and tensor mappings for modern bert support, model not supported yet but working on getting conversion to work for encoder only	2025-08-21 12:38:04 -04:00
David Zhao	79c1160b07	cuda: refactored ssm_scan and use CUB (#13291 ) * cuda: refactored ssm_scan to use CUB * fixed compilation error when when not using CUB * assign L to constant and use size_t instead of int * deduplicated functions * change min blocks per mp to 1 * Use cub load and store warp transpose * suppress clang warning	2025-08-09 20:29:43 +02:00
Aman Gupta	34c9d765bf	CUDA: add attention sinks for tile and wmma (#15178 ) * CUDA: add attention sinks for tile and wmma * Review: formatting changes + remove syncthreads from tile + remove warp_reduce_max from wmma	2025-08-09 20:00:24 +08:00
compilade	e54d41befc	gguf-py : add Numpy MXFP4 de/quantization support (#15111 ) * gguf-py : add MXFP4 de/quantization support * ggml-quants : handle zero amax for MXFP4	2025-08-08 17:48:26 -04:00
Johannes Gäßler	4850b52aed	server-bench: external OAI servers, sqlite (#15179 ) * server-bench: external OAI servers, sqlite * Update scripts/server-bench.py Co-authored-by: Sigbjørn Skjæret <sigbjorn.skjaeret@scala.com> * Update scripts/server-bench.py Co-authored-by: Sigbjørn Skjæret <sigbjorn.skjaeret@scala.com> * Update scripts/server-bench.py Co-authored-by: Sigbjørn Skjæret <sigbjorn.skjaeret@scala.com> * raise_for_status --------- Co-authored-by: Sigbjørn Skjæret <sigbjorn.skjaeret@scala.com>	2025-08-08 23:04:36 +02:00
AN Long	cd6983d56d	ggml : fix field name when new ggml_backend (#14944 )	2025-08-08 14:37:22 +02:00
Olivier Chafik	6c7e9a5440	vendor: sync minja (#15161 ) * vendor: sync minja * Update minja.hpp * Apply suggestions from code review Co-authored-by: Sigbjørn Skjæret <sigbjorn.skjaeret@scala.com> --------- Co-authored-by: Sigbjørn Skjæret <sigbjorn.skjaeret@scala.com>	2025-08-08 10:45:18 +01:00
Johannes Gäßler	1425f587a8	CUDA: attention sinks for mma FlashAttention (#15157 )	2025-08-08 08:19:58 +02:00
lhez	aaa3d07ae7	opencl: support sink in `soft_max` (attn sinks) (#15152 )	2025-08-07 21:47:03 -07:00
Xuan-Son Nguyen	50aa938901	convert : support non-mxfp4 HF model (#15153 ) * convert : support non-mxfp4 HF model * rm redundant check * disable debug check	2025-08-07 23:26:03 +02:00
Jeff Bolz	c4f53563df	vulkan: support fattn sinks (#15126 )	2025-08-07 22:44:20 +02:00
Jeff Bolz	a0552c8bee	vulkan: Add env var to disable host visible vidmem (#15109 )	2025-08-07 22:07:11 +02:00
RunningLeon	99acbc9921	llama : Support intern-s1 (#14875 ) * support internvl * support interns1 * resolve comments * put interns1 in tensor mapping * resolve comment * move tokenizer changes to sub class	2025-08-07 18:20:40 +02:00
uvos	7ad67ba9fe	HIP: add cmake option to enable compiler output of kernel resource usage metrics (#15103 )	2025-08-07 16:44:14 +02:00
Christian Kastner	9a96389544	ggml: Skip backend library linking code when GGML_BACKEND_DL=ON (#15094 ) Any available libraries are found and loaded dynamically at runtime.	2025-08-07 13:45:41 +02:00
Johannes Gäßler	1d72c84188	CUDA: GEMM for FP32/FP16/BF16 and ne11 <= 16 (#15131 ) * CUDA: GEMM for FP32/FP16/BF16 and ne11 <= 16	2025-08-07 10:53:21 +02:00
Johannes Gäßler	20638e4f16	scripts: fix crash when --tool is not set (#15133 )	2025-08-07 08:50:30 +02:00
Daniel Bevenius	36d3f00e14	requirements : fix PyTorch uint64 compatibility (#15134 ) This commit addresses an issue with the convert_hf_to_gguf script which is currently failing with: ```console AttributeError: module 'torch' has no attribute 'uint64' ``` This occurred because safetensors expects torch.uint64 to be available in the public API, but PyTorch 2.2.x only provides limited support for unsigned types beyond uint8 it seems. The torch.uint64 dtype exists but is not exposed in the standard torch namespace (see pytorch/pytorch#58734). PyTorch 2.4.0 properly exposes torch.uint64 in the public API, resolving the compatibility issue with safetensors. This also required torchvision to updated to =0.19.0 for compatibility. Refs: https://huggingface.co/spaces/ggml-org/gguf-my-repo/discussions/186#68938de803e47d990aa087fb Refs: https://github.com/pytorch/pytorch/issues/58734	2025-08-07 05:31:48 +02:00
Reese Levine	5fd160bbd9	ggml: Add basic SET_ROWS support in WebGPU (#15137 ) * Begin work on set_rows * Work on set rows * Add error buffers for reporting unsupported SET_ROWS indices * Remove extra comments	2025-08-06 15:14:40 -07:00
rmatif	756cfea826	fix profiling crash (#15072 )	2025-08-06 14:17:51 -07:00
lhez	e725a1a982	opencl: add `swiglu_oai` and `add_id` (#15121 ) * opencl: add `swiglu-oai` * opencl: add `add_id` * opencl: add missing `add_id.cl`	2025-08-06 12:12:17 -07:00
Sachin Desai	3db4da56a5	chat : support Granite model reasoning and tool call (#14864 )	2025-08-06 20:27:30 +02:00
Juk Armstrong	476aa3fd57	Fixed name `-override-tensors` to `-override-tensor` (#15129 )	2025-08-06 17:28:48 +01:00
Diego Devesa	0d8831543c	ggml : fix fallback to CPU for ununsupported ops (#15118 )	2025-08-06 14:37:35 +02:00
Sigbjørn Skjæret	65c797c4fa	chat : fix yandex chat template (#15116 )	2025-08-06 13:26:49 +02:00
stevenkuang	25726898e8	chat : fix hunyuan auto-detection (#15114 ) Signed-off-by: stevenkuang <stevenkuang@tencent.com>	2025-08-06 11:48:30 +02:00
Chenguang Li	2241453252	CANN: add support for ACL Graph (#15065 ) * feat(cann): add optional support for ACL Graph execution This commit adds support for executing ggml computational graphs using Huawei's ACL graph mode via the USE_CANN_GRAPH flag. The support can be enabled at compile time using the CMake option: -DUSE_CANN_GRAPH=ON By default, ACL graph execution is disabled, and the fallback path uses node-by-node execution. Key additions: - CMake option to toggle graph mode - Graph capture and execution logic using - Tensor property matching to determine whether graph update is required - Safe fallback and logging if the environment variable LLAMA_SET_ROWS is unset or invalid This prepares the backend for performance improvements in repetitive graph execution scenarios on Ascend devices. Signed-off-by: noemotiovon <757486878@qq.com> * Fix review comments Signed-off-by: noemotiovon <757486878@qq.com> * remane USE_CANN_GRAPH to USE_ACL_GRAPH Signed-off-by: noemotiovon <757486878@qq.com> * fix typo Signed-off-by: noemotiovon <757486878@qq.com> --------- Signed-off-by: noemotiovon <757486878@qq.com>	2025-08-06 14:12:42 +08:00
Reese Levine	9515c6131a	ggml: WebGPU disable SET_ROWS for now (#15078 ) * Add paramater buffer pool, batching of submissions, refactor command building/submission * Add header for linux builds * Free staged parameter buffers at once * Format with clang-format * Fix thread-safe implementation * Use device implicit synchronization * Update workflow to use custom release * Remove testing branch workflow * Disable set_rows until it's implemented * Fix potential issue around empty queue submission * Try synchronous submission * Try waiting on all futures explicitly * Add debug * Add more debug messages * Work on getting ssh access for debugging * Debug on failure * Disable other tests * Remove extra if * Try more locking * maybe passes? * test * Some cleanups * Restore build file * Remove extra testing branch ci	2025-08-05 16:26:38 -07:00
Georgi Gerganov	fd1234cb46	llama : add gpt-oss (#15091 ) * oai moe * compat with new checkpoint * add attn sink impl * add rope scaling yarn * logits match with latest transformers code * wip chat template * rm trailing space * use ggml_scale_bias * rm redundant is_swa_all * convert interleaved gate_up * graph : fix activation function to match reference (#7) * vocab : handle o200k_harmony special tokens * ggml : add attention sinks support (#1) * llama : add attn sinks * ggml : add attn sinks * cuda : add attn sinks * vulkan : add support for sinks in softmax remove unnecessary return * ggml : add fused swiglu_oai op (#11) * ggml : add fused swiglu_oai op * Update ggml/src/ggml-cpu/ops.cpp Co-authored-by: Georgi Gerganov <ggerganov@gmail.com> * update CUDA impl * cont : metal impl * add vulkan impl * test-backend-ops : more test cases, clean up * llama : remove unfused impl * remove extra lines --------- Co-authored-by: Georgi Gerganov <ggerganov@gmail.com> --------- Co-authored-by: slaren <slarengh@gmail.com> * repack mxfp4 upon conversion * clean up a bit * enable thinking * add quick hack to render only some special tokens * fix bf16 conversion * remove vocab hack * webui ok * support chat parsing for gpt-oss * fix webui * direct mapping mxfp4, FINALLY * force using mxfp4 * properly use lazy tensor * ggml : add mxfp4 ggml : use e8m0 conversion instead of powf Co-authored-by: Diego Devesa <slarengh@gmail.com> change kvalues_mxfp4 table to match e2m1 (#6) metal : remove quantization for now (not used) cuda : fix disabled CUDA graphs due to ffn moe bias vulkan : add support for mxfp4 cont : add cm2 dequant * ggml : add ggml_add_id (#13) * ggml : add ggml_add_id * add cuda impl * llama : add weight support check for add_id * perf opt * add vulkan impl * rename cuda files * add metal impl * allow in-place ggml_add_id * llama : keep biases on CPU with --cpu-moe * llama : fix compile error ggml-ci * cuda : add fallback for __nv_cvt_e8m0_to_bf16raw ggml-ci * cleanup ggml-ci * sycl : fix supports_op for MXFP4 ggml-ci * fix Unknown reasoning format * ggml-cpu : fix AVX build ggml-ci * fix hip build ggml-ci * cuda : add mxfp4 dequantization support for cuBLAS ggml-ci * ggml-cpu : fix mxfp4 fallback definitions for some architectures ggml-ci * cuda : fix version required for __nv_cvt_e8m0_to_bf16raw --------- Co-authored-by: Xuan Son Nguyen <son@huggingface.co> Co-authored-by: slaren <slarengh@gmail.com>	2025-08-05 22:10:36 +03:00
Sigbjørn Skjæret	f324a3b715	chat : only remove double bos/eos if added (#15086 ) * only remove double bos/eos if added * fix tests	2025-08-05 20:43:36 +02:00
Georgi Gerganov	be42642581	readme : update hot topics (#15097 )	2025-08-05 20:19:33 +03:00
Romain Biessy	3306ceabf0	sycl: fix mul_mat selection (#15092 )	2025-08-05 18:39:55 +02:00
Juk Armstrong	c81de6e107	Fix `glm4moe` bug (#15088 )	2025-08-05 13:56:44 +01:00
Alex Wu	22f060c9c4	webui: fix markdown table (#15081 ) * webui: fix markdown table * webui: fix table display with themes	2025-08-05 13:56:44 +02:00

1 2 3 4 5 ...

6140 Commits All Branches Search

6140 Commits

All Branches