llama.cpp

History

Adrien Gallouët 4e76d24f28 ggml : fix AMX and add batched support (#19925 ) llama-perplexity -hf ggml-org/Qwen3-0.6B-GGUF:Q4_0 -f wikitext-2-raw/wiki.test.raw -c 2048 -b 2048 --chunks 2 before this commit: ``` perplexity: calculating perplexity over 2 chunks, n_ctx=2048, batch_size=2048, n_seq=1 perplexity: 2.31 seconds per pass - ETA 0.07 minutes [1]17.3868,[2]22.2199, Final estimate: PPL = 22.2199 +/- 1.59692 llama_perf_context_print: load time = 878.56 ms llama_perf_context_print: prompt eval time = 2037.82 ms / 4096 tokens ( 0.50 ms per token, 2009.99 tokens per second) llama_perf_context_print: eval time = 0.00 ms / 1 runs ( 0.00 ms per token, inf tokens per second) llama_perf_context_print: total time = 6403.17 ms / 4097 tokens llama_perf_context_print: graphs reused = 0 llama_memory_breakdown_print: \| memory breakdown [MiB] \| total free self model context compute unaccounted \| llama_memory_breakdown_print: \| - Host \| 845 = 318 + 224 + 302 \| llama_memory_breakdown_print: \| - CPU_REPACK \| 288 = 288 + 0 + 0 \| llama_memory_breakdown_print: \| - AMX \| 31 = 31 + 0 + 0 \| ``` after this commit: ``` perplexity: calculating perplexity over 2 chunks, n_ctx=2048, batch_size=2048, n_seq=1 perplexity: 1.98 seconds per pass - ETA 0.05 minutes [1]17.2005,[2]21.8220, Final estimate: PPL = 21.8220 +/- 1.56485 llama_perf_context_print: load time = 719.23 ms llama_perf_context_print: prompt eval time = 1676.23 ms / 4096 tokens ( 0.41 ms per token, 2443.58 tokens per second) llama_perf_context_print: eval time = 0.00 ms / 1 runs ( 0.00 ms per token, inf tokens per second) llama_perf_context_print: total time = 4258.74 ms / 4097 tokens llama_perf_context_print: graphs reused = 0 llama_memory_breakdown_print: \| memory breakdown [MiB] \| total free self model context compute unaccounted \| llama_memory_breakdown_print: \| - Host \| 845 = 318 + 224 + 302 \| llama_memory_breakdown_print: \| - AMX \| 319 = 319 + 0 + 0 \| ``` (no more CPU_REPACK) after this commit, disabling amx: ``` perplexity: calculating perplexity over 2 chunks, n_ctx=2048, batch_size=2048, n_seq=1 perplexity: 2.34 seconds per pass - ETA 0.07 minutes [1]17.2005,[2]21.8220, Final estimate: PPL = 21.8220 +/- 1.56485 llama_perf_context_print: load time = 841.91 ms llama_perf_context_print: prompt eval time = 2057.28 ms / 4096 tokens ( 0.50 ms per token, 1990.98 tokens per second) llama_perf_context_print: eval time = 0.00 ms / 1 runs ( 0.00 ms per token, inf tokens per second) llama_perf_context_print: total time = 6454.51 ms / 4097 tokens llama_perf_context_print: graphs reused = 0 llama_memory_breakdown_print: \| memory breakdown [MiB] \| total free self model context compute unaccounted \| llama_memory_breakdown_print: \| - Host \| 845 = 318 + 224 + 302 \| llama_memory_breakdown_print: \| - CPU_REPACK \| 319 = 319 + 0 + 0 \| ``` => same perplexity. Signed-off-by: Adrien Gallouët <angt@huggingface.co>		2026-02-26 21:39:11 +01:00
..
ggml-blas	ggml : add ggml_build_forward_select (#18550 )	2026-01-19 20:03:19 +02:00
ggml-cann	CANN: Remove unnecessary wrapper for `gml_backend_buft_is_cann` (#18968 )	2026-02-10 14:19:30 +08:00
ggml-cpu	ggml : fix AMX and add batched support (#19925 )	2026-02-26 21:39:11 +01:00
ggml-cuda	Improve CUDA graph capture (#19754 )	2026-02-21 15:09:36 +05:30
ggml-hexagon	hexagon refactor all Ops to use local context struct (#19819 )	2026-02-23 16:32:14 -08:00
ggml-hip	HIP: add mmf for CDNA (#18896 )	2026-01-29 11:10:53 +01:00
ggml-metal	models : optimize qwen3next graph (#19375 )	2026-02-14 12:57:36 +02:00
ggml-musa	CUDA: faster tile FA, add oob checks, more HSs (#16492 )	2025-10-11 20:54:32 +02:00
ggml-opencl	opencl: refactor expm1 and softplus (#19404 )	2026-02-17 14:47:18 -08:00
ggml-rpc	rpc : use unordered_map::reserve and emplace (#18513 )	2026-01-02 12:09:36 +02:00
ggml-sycl	support permuted, remove check s0/s10 (#19889 )	2026-02-26 10:27:20 +08:00
ggml-virtgpu	ggml-virtgpu: improve the reliability of the code (#19846 )	2026-02-26 20:00:57 +08:00
ggml-vulkan	vulkan: fix fp16 Flash Attention on Windows AMD RDNA2 and below (#19921 )	2026-02-26 19:11:04 +01:00
ggml-webgpu	ggml-webgpu: Add unary op (SQR, SQRT, SIN, COS) support. (#19700 )	2026-02-19 09:18:30 -07:00
ggml-zdnn	ggml-zdnn : mark zDNN buffers as non-host (#18967 )	2026-01-22 01:16:21 +01:00
ggml-zendnn	ggml-zendnn : resolve ZenDNN backend cross-module symbol dependency (#19159 )	2026-01-29 12:28:57 +08:00
CMakeLists.txt	hexagon: enable offloading to Hexagon on Windows on Snapdragon (#19150 )	2026-01-29 12:33:21 -08:00
ggml-alloc.c	ggml : make `ggml_is_view` as API (#19539 )	2026-02-16 17:43:34 +02:00
ggml-backend-dl.cpp	hexagon: enable offloading to Hexagon on Windows on Snapdragon (#19150 )	2026-01-29 12:33:21 -08:00
ggml-backend-dl.h	hexagon: enable offloading to Hexagon on Windows on Snapdragon (#19150 )	2026-01-29 12:33:21 -08:00
ggml-backend-impl.h	llama: use host memory if device reports 0 memory (#18587 )	2026-01-09 05:34:56 +08:00
ggml-backend-reg.cpp	ggml : use noexcept overload for is_regular_file in backend registration (#19452 )	2026-02-10 10:57:48 +01:00
ggml-backend.cpp	ggml-backend: fix async set/get fallback sync (#19179 )	2026-02-02 10:00:05 +01:00
ggml-common.h	llama : add gpt-oss (#15091 )	2025-08-05 22:10:36 +03:00
ggml-impl.h	ggml : make `ggml_is_view` as API (#19539 )	2026-02-16 17:43:34 +02:00
ggml-opt.cpp	finetune: SGD optimizer, more CLI args (#13873 )	2025-08-14 12:03:57 +02:00
ggml-quants.c	ggml : fix uninitialized is_on_grid in quantize_row_iq3_xxs_impl (#15928 )	2025-09-23 10:25:20 +02:00
ggml-quants.h	llama : add gpt-oss (#15091 )	2025-08-05 22:10:36 +03:00
ggml-threading.cpp	…
ggml-threading.h	…
ggml.c	ggml/gguf : prevent integer overflows (#19856 )	2026-02-24 20:17:11 +02:00
ggml.cpp	…
gguf.cpp	gguf : avoid too many file size calls (#19919 )	2026-02-26 12:46:32 +02:00