llama.cpp

Commit Graph

Author	SHA1	Message	Date
ddh0	2a3f579d1f	does this fix it?	2025-12-14 01:55:02 -06:00
ddh0	9613c48172	with logging	2025-12-14 00:36:59 -06:00
ddh0	d1e5c60442	add missing values to `common_params_sampling::print()`	2025-12-13 23:26:03 -06:00
ddh0	965bcc9dc4	fix leftover `window_size`	2025-12-13 22:19:15 -06:00
ddh0	b8a9626a73	oops forgot args.cpp	2025-12-13 22:17:08 -06:00
ddh0	a96ddd743a	re-write + change parameters + simplify	2025-12-13 22:15:03 -06:00
ddh0	67a733670e	Merge branch 'ggml-org:master' into power-law-sampler	2025-12-13 17:27:35 -06:00
Xuan-Son Nguyen	c00ff929dc	scripts: add script to compare logprobs of llama.cpp against other frameworks (#17947 ) * scripts: add script to compare logits of llama.cpp against other frameworks * accept custom prompt file * fix code style * clarify endpoint * fix displaying * use abs for diff * fix vllm case * rm output file * rename to compare-logprobs * add "pattern"	2025-12-13 22:33:29 +01:00
Sergey Fedorov	4ed2bae50d	server-models.cpp: add missing <filesystem> (#18000 ) Fixes: https://github.com/ggml-org/llama.cpp/issues/17999	2025-12-13 22:02:43 +01:00
Jeff Bolz	5266379bca	llama_context: synchronize before reallocating output buffer (#17974 )	2025-12-13 09:19:51 -06:00
Xuan-Son Nguyen	4d5ae24c0a	arg: fix common_params_parse not accepting negated arg (#17991 )	2025-12-13 12:53:37 +01:00
Gustavo Rocha Dias	66ba51252e	cmake: correct scope - link ws2_32 for MinGW/w64devkit builds in cpp-httplib (#17972 ) * fix - w64devkit build * fix - w64devkit build private scope	2025-12-13 12:46:36 +01:00
Jeff Bolz	36255a2268	vulkan: support get_rows for i32 (#17941 )	2025-12-13 10:12:53 +01:00
Jeff Bolz	3229a23fa6	vulkan: support GGML_OP_DIAG (#17893 )	2025-12-13 10:07:49 +01:00
Jeff Bolz	303f8615e9	vulkan: Multi-pass softmax for large number of cols (#17892 ) When the number of cols is large, split each row across multiple workgroups. There are three phases that communicate partial results through temp buffers: (1) compute max partials (2) take max of partials, compute sum(exp(x-max)) partials (3) sum partials, compute scaled result	2025-12-13 10:04:29 +01:00
Georgi Gerganov	3c6391e748	speculative-simple : free batch on exit (#17985 )	2025-12-13 09:48:34 +02:00
Sigbjørn Skjæret	8e4d678528	common : skip model validation when --completion-bash is requested (#17975 )	2025-12-13 08:40:50 +01:00
Jeff Bolz	07a10c1090	vulkan: Allow non-pow2 n_experts in topk_moe (#17872 )	2025-12-13 08:40:04 +01:00
Sigbjørn Skjæret	2bc94e7928	add llama-completion to completion-bash executables (#17976 )	2025-12-13 08:35:50 +01:00
Daniel Bevenius	fd1085ffb7	model-conversion : use CONVERTED_MODEL value for converted model [no ci] (#17984 ) * model-conversion : use CONVERTED_MODEL value for converted model [no ci] This commit updates the model verification scripts to use the CONVERTED_MODEL environment variable instead of using the MODEL_PATH (the original model path) as the basis for the converted model file name. The motivation for this that currently if the converted model file name differs from the original model directory/name the verification scripts will look for the wrong .bin files that were generating when running the models. For example, the following steps were not possible: ```console (venv) $ huggingface-cli download google/gemma-3-270m-it --local-dir ggml-org/gemma-3-270m (venv) $ python3 convert_hf_to_gguf.py ggml-org/gemma-3-270m --outfile test-bf16.gguf --outtype bf16 (venv) $ cd examples/model-conversion/ (venv) $ export MODEL_PATH=../../ggml-org/gemma-3-270m (venv) $ export CONVERTED_MODEL=../../test-bf16.gguf (venv) $ make causal-verify-logits ... Data saved to data/llamacpp-test-bf16.bin Data saved to data/llamacpp-test-bf16.txt Error: llama.cpp logits file not found: data/llamacpp-gemma-3-270m.bin Please run scripts/run-converted-model.sh first to generate this file. make: *** [Makefile:62: causal-verify-logits] Error 1 ``` With the changes in this commit, the above steps will now work as expected.	2025-12-13 08:34:26 +01:00
ddh0	1879fc6dc6	Merge branch 'ggml-org:master' into power-law-sampler	2025-12-13 01:17:53 -06:00
ddh0	824bb3aa6e	fix compiler warning, add commented-out logging per token	2025-12-13 00:23:15 -06:00
ddh0	0a19a3fd6c	remove old debug log, style nit	2025-12-12 23:45:45 -06:00
ddh0	94cb883ed9	copy from author ref: https://gist.github.com/MrJackSpade/9be99c7efbba7b95a41377e123b7b069	2025-12-12 23:19:08 -06:00
ddh0	53380c183f	add missing parameters in `server-task.cpp`	2025-12-12 22:39:51 -06:00
Xuan-Son Nguyen	380b4c984e	common: support negated args (#17919 ) * args: support negated args * update docs * fix typo * add more neg options * Apply suggestions from code review Co-authored-by: Sigbjørn Skjæret <sigbjorn.skjaeret@scala.com> * rm duplicated arg * fix LLAMA_ARG_NO_HOST * add test --------- Co-authored-by: Sigbjørn Skjæret <sigbjorn.skjaeret@scala.com>	2025-12-12 23:58:53 +01:00
Xuan-Son Nguyen	e39a2ce66d	clip: move model cgraphs into their own files (#17965 ) * clip: move model cgraphs into their own files * more explicit enums * fix linux build * fix naming * missing headers * nits: add comments for contributors	2025-12-12 21:14:48 +01:00
jiahao su	a8c7f33d79	ci : change the cann version and the container pull method (#17953 ) fix error format Update build.yml Remove unnecessary zip files fix update	2025-12-12 20:43:00 +01:00
Sigbjørn Skjæret	b7f5f46e03	docker : include legacy llama-completion binary (#17964 )	2025-12-12 19:39:23 +01:00
Johannes Gäßler	482211438d	CUDA: fix overflow in MMA kernel without stream-k (#17939 )	2025-12-12 17:43:58 +01:00
Georgi Gerganov	7bed317f53	models : fix the attn_factor for mistral3 graphs + improve consistency (#17945 ) * models : fix the attn_factor for mistral3 graphs * cont : rework attn_factor correction logic * cont : make deepseek2 consistent * cont : add TODO * cont : special-case DSv2 * cont : revert Mistral 3 Large changes * cont : fix DS2 to use the original attn_factor * cont : minor comments	2025-12-12 17:12:40 +02:00
Sigbjørn Skjæret	dcb7d17758	cann : fix ops broken by circular padding guard (#17825 )	2025-12-12 15:49:27 +01:00
ixgbe	51604435e8	ggml-cpu : fix RISC-V Q4_0 repack select and RVV feature reporting (#17951 ) * ggml-cpu:fix RISC-V Q4_0 repack select and RVV feature reporting Signed-off-by: Wang Yang <yangwang@iscas.ac.cn> * using the name VLEN instead of CNT * Update ggml/include/ggml-cpu.h --------- Signed-off-by: Wang Yang <yangwang@iscas.ac.cn> Co-authored-by: Georgi Gerganov <ggerganov@gmail.com>	2025-12-12 16:26:03 +02:00
Xuan-Son Nguyen	17158965ac	mtmd: explicitly forbidden inclusion of private header and libcommon (#17946 )	2025-12-12 15:16:06 +01:00
Aleksander Grygier	12280ae905	webui: Fix parsing non-LaTeX occurrencies of `$` or `$` (#17810 ) * fix: Improve latex protection logic to prevent turning non-latex `\(` into `$` * chore: update webui build output	2025-12-12 15:13:36 +01:00
Xuan-Son Nguyen	54a0fee4b7	arg: add -mm and -mmu as short form of --mmproj and --mmproj-url (#17958 ) * arg: add -mm and -mmu as short form of --mmproj and --mmproj-url * correct order * update docs	2025-12-12 14:06:06 +01:00
Daniel Bevenius	dada4c846d	model-conversion : remove max diff check in compare-logits [no ci] (#17954 ) This commit removes the maximum difference check from the compare-logits.py which would stop early if the difference between the logits exceeded a threshold. The motivation for removing this is that it can be useful to be able to get the complete log for debugging/reporting purposes.	2025-12-12 13:25:16 +01:00
Adrien Gallouët	b8ee22cfde	common : add minimalist multi-thread progress bar (#17602 ) Signed-off-by: Adrien Gallouët <angt@huggingface.co>	2025-12-12 12:44:35 +01:00
Gustavo Rocha Dias	2eaa2c65cb	cmake: link ws2_32 for MinGW/w64devkit builds in cpp-httplib (#17949 )	2025-12-12 12:02:28 +01:00
yulo	c33a58bced	HIP: enable mmf for RDNA3 (#17879 ) * enable mmf for RDNA3 * disable mmf for some shape * move some mmvf to mmf * more mmfv to mmf * 3 is good in mmvf --------- Co-authored-by: zhang hui <you@example.com>	2025-12-12 11:34:33 +01:00
ddh0	5c78b7927f	oops, straggler	2025-12-11 22:47:36 -06:00
ddh0	2d62bbea9f	remove `target_range` param, make `target == 1` no-op, cleanup code	2025-12-11 22:43:10 -06:00
ddh0	dcada035b4	add missing enums	2025-12-11 17:49:47 -06:00
ddh0	534cb4fbba	clarify behaviour when `window_size = 0`	2025-12-11 17:29:04 -06:00
ddh0	cd7de7c7a8	add power law case to `common_sampler_init`, add sampler name mappings	2025-12-11 17:23:27 -06:00
ddh0	b3aea57768	minor	2025-12-11 16:48:52 -06:00
ddh0	93169593b8	remove old unused code from algorithm	2025-12-11 16:46:17 -06:00
ddh0	f3457a83e6	minor	2025-12-11 16:36:00 -06:00
ddh0	4959878a74	improved comments	2025-12-11 16:27:14 -06:00
ddh0	ffe163911b	add args, rename `queue_size` -> `window_size`	2025-12-11 15:16:11 -06:00

1 2 3 4 5 ...

7416 Commits All Branches Search

7416 Commits

All Branches