llama.cpp

History

Daniel Bevenius 4150da9a95 examples : add --kv-unified to batched example (#18774 ) This commit adds the --kv-unified flag to the batched example. This flag is currently specified in the README.md as required, but is currently not available as a command line option for the batched example. The motivation for this is that specifying this flag as the README instructs, will lead to an error about the flag not being recognized, and without this option the example fail with the following error: ```console split_equal: sequential split is not supported when there are coupled sequences in the input batch (you may need to use the -kvu flag) decode: failed to find a memory slot for batch of size 4 main: llama_decode() failed ```		2026-01-12 13:47:58 +01:00
..
batched	examples : add --kv-unified to batched example (#18774 )	2026-01-12 13:47:58 +01:00
batched.swift	examples : remove references to `make` in examples [no ci] (#15457 )	2025-08-21 06:12:28 +02:00
convert-llama2c-to-ggml	gguf: gguf_writer refactor (#15691 )	2025-09-05 11:34:28 +02:00
debug	debug : include LLAMA_POOLING_TYPE_UNSPECIFIED in pooling check (#18692 )	2026-01-11 16:34:41 +01:00
deprecation-warning	Update deprecation-warning.cpp (#10619 )	2024-12-04 23:19:20 +01:00
diffusion	llama : add `use_direct_io` flag for model loading (#18166 )	2026-01-08 08:35:30 +02:00
embedding	model : add LFM2-ColBert-350M (#18607 )	2026-01-05 19:52:56 +01:00
eval-callback	common : refactor common_sampler + grammar logic changes (#17937 )	2025-12-14 10:11:13 +02:00
gen-docs	gen-docs: automatically update markdown file (#18294 )	2025-12-22 19:30:19 +01:00
gguf	examples(gguf): GGUF example outputs (#17025 )	2025-11-05 19:58:16 +02:00
gguf-hash	GGUF: C++ refactor, backend support, misc fixes (#11030 )	2025-01-07 18:01:58 +01:00
idle	metal : add residency sets keep-alive heartbeat (#17766 )	2025-12-05 19:38:54 +02:00
llama.android	android: routine maintenance - Dec 2025 (#18338 )	2025-12-29 15:51:13 +02:00
llama.swiftui	llama : deprecate llama_kv_self_ API (#14030 )	2025-06-06 14:11:15 +03:00
lookahead	common : refactor common_sampler + grammar logic changes (#17937 )	2025-12-14 10:11:13 +02:00
lookup	common : refactor common_sampler + grammar logic changes (#17937 )	2025-12-14 10:11:13 +02:00
model-conversion	model-conversion : add warn about transformers mismatch (#18691 )	2026-01-08 09:29:53 +01:00
parallel	common : refactor common_sampler + grammar logic changes (#17937 )	2025-12-14 10:11:13 +02:00
passkey	examples : remove references to `make` in examples [no ci] (#15457 )	2025-08-21 06:12:28 +02:00
retrieval	model : add LFM2-ColBert-350M (#18607 )	2026-01-05 19:52:56 +01:00
save-load-state	common : refactor common_sampler + grammar logic changes (#17937 )	2025-12-14 10:11:13 +02:00
simple	examples : support encoder-decoder models in the simple example (#16002 )	2025-09-17 10:29:00 +03:00
simple-chat	simple-chat : fix context-exceeded condition (#14494 )	2025-07-02 14:12:07 +03:00
simple-cmake-pkg	examples : add missing code block end marker [no ci] (#17756 )	2025-12-04 14:17:30 +01:00
speculative	common : restore grammar-based rejection sampling (#18137 )	2025-12-17 19:46:00 +02:00
speculative-simple	common : refactor common_sampler + grammar logic changes (#17937 )	2025-12-14 10:11:13 +02:00
sycl	[SYCL] replace llama-cli by llama-completion to rm the impact to test script (#18290 )	2025-12-23 12:59:12 +08:00
training	common : refactor common_sampler + grammar logic changes (#17937 )	2025-12-14 10:11:13 +02:00
CMakeLists.txt	examples : add debug utility/example (#18464 )	2026-01-07 10:42:19 +01:00
convert_legacy_llama.py	metadata: Detailed Dataset Authorship Metadata (#8875 )	2024-11-13 21:10:38 +11:00
json_schema_pydantic_example.py	py : type-check all Python scripts with Pyright (#8341 )	2024-07-07 15:04:39 -04:00
json_schema_to_grammar.py	common : fix json schema with '\' in literals (#17307 )	2025-11-29 17:06:32 +01:00
llama.vim	llama : remove KV cache defragmentation logic (#15473 )	2025-08-22 12:22:13 +03:00
pydantic_models_to_grammar.py	pydantic : replace uses of __annotations__ with get_type_hints (#8474 )	2024-07-14 19:51:21 -04:00
pydantic_models_to_grammar_examples.py	llama : move end-user examples to tools directory (#13249 )	2025-05-02 20:27:13 +02:00
reason-act.sh	scripts : make the shell scripts cross-platform (#14341 )	2025-06-30 10:17:18 +02:00
regex_to_grammar.py	py : switch to snake_case (#8305 )	2024-07-05 07:53:33 +03:00
server-llama2-13B.sh	scripts : make the shell scripts cross-platform (#14341 )	2025-06-30 10:17:18 +02:00
server_embd.py	llama : fix FA when KV cache is not used (i.e. embeddings) (#12825 )	2025-04-08 19:54:51 +03:00
ts-type-to-grammar.sh	scripts : make the shell scripts cross-platform (#14341 )	2025-06-30 10:17:18 +02:00