llama.cpp

History

Andrea Arcangeli 990e4d9698 common/grammar: fix grammar parsing issues to prevent stack overflow and hangs (#18604 ) * grammar: add test case for nullable symbol loop Reproduce stack overflow (or OOM) with ( [x]* )* found while adding GBNF support to ripgrep-edit. llama-server reproducer: curl \ -X POST \ -d '{ "messages": [{ "role": "user", "content": "write yes" }], "grammar": "root ::= ( [x]* )" }' \ -H "Content-Type: application/json" \ http://localhost:8811/v1/chat/completions grammar: prevent stack overflow with nullable symbol loop Fix a potential stack overflow in llama_grammar_advance_stack that could occur when processing grammars with nullable symbols that lead to infinite derivations of empty strings. The fix introduces cycle detection by tracking visited stacks to prevent infinite recursion. rg-edit regexp: llama_grammar_advance_stack rg-edit extra-args: -A20 rg-edit directive: """Rewrite: fix the following segfault: [..] ⚫ Testing segfault. Grammar: root ::= ( [x]* )* root ::= ( [x]* )* Segmentation fault build/bin/test-grammar-integration""" gptel-context: (("~/llama.cpp/src/llama-grammar.cpp") ("~/llama.cpp/tests/test-grammar-integration.cpp") ("~/llama.cpp/grammars/./list.gbnf") ("~/llama.cpp/grammars/./json_arr.gbnf") ("~/llama.cpp/grammars/./json.gbnf") ("~/llama.cpp/grammars/./japanese.gbnf") ("~/llama.cpp/grammars/./english.gbnf") ("~/llama.cpp/grammars/./chess.gbnf") ("~/llama.cpp/grammars/./c.gbnf") ("~/llama.cpp/grammars/./arithmetic.gbnf") ("~/llama.cpp/grammars/./README.md")) * grammar: convert recursive llama_grammar_advance_stack to iterative This change converts the function to an iterative approach using explicit stacks, which prevents deep recursion and eliminates the risk of stack overflow. rg-edit regexp: llama_grammar_advance_stack rg-edit extra-args: -A30 rg-edit directive: """Rewrite: fix the following segfault: [..] ⚫ Testing segfault. Grammar: root ::= ( [x]* )* root ::= ( [x]* )* Segmentation fault build/bin/test-grammar-integration convert from recursive to interactive""" gptel-context: (("~/llama.cpp/src/llama-grammar.cpp") ("~/llama.cpp/tests/test-grammar-integration.cpp") ("~/llama.cpp/grammars/./list.gbnf") ("~/llama.cpp/grammars/./json_arr.gbnf") ("~/llama.cpp/grammars/./json.gbnf") ("~/llama.cpp/grammars/./japanese.gbnf") ("~/llama.cpp/grammars/./english.gbnf") ("~/llama.cpp/grammars/./chess.gbnf") ("~/llama.cpp/grammars/./c.gbnf") ("~/llama.cpp/grammars/./arithmetic.gbnf") ("~/llama.cpp/grammars/./README.md")) v2: Added a `std::set` to perform tree-based lookups with O(N log N) complexity. Testing with a parallel run of `test-grammar-integration` shows a double-digit percentage increase in runtime. An `unordered_set` with O(1) hashing was also evaluated, but the overhead of constructing hash keys from pointers made it significantly slower than the rbtree implementation that only requires an ordering operator. The performance regression in the test suite appears justified by the overall reduction in algorithmic complexity. Co-developed-by: Piotr Wilkin (ilintar) <piotr.wilkin@syndatis.com> * grammar: add test case for hang in repetition grammar processing This commit adds a new test case to the grammar integration tests that specifically targets a hang scenario in the repetition grammar parser found while adding GBNF support to ripgrep-edit. llama-server reproducer: curl \ -X POST \ -d '{ "messages": [{ "role": "user", "content": "write yes" }], "grammar": "root ::= (([^x]){0,99}){0,99}" }' \ -H "Content-Type: application/json" \ http://localhost:8811/v1/chat/completions grammar: add repetition threshold check The change introduces a maximum repetition threshold to avoid excessive rule expansion during grammar parsing. When parsing repetition patterns like {m,n}, the parser now calculates the potential number of rules that would be generated and throws an error if the product of previous rules and new rules exceeds the threshold. A test case was added to verify the threshold is properly enforced for deeply nested repetition patterns that would otherwise cause hangs.		2026-03-21 18:43:35 +01:00
..
models	model : add control vector support where missing (#20653 )	2026-03-18 23:25:12 +01:00
CMakeLists.txt	model : add Jina Embeddings v5 Nano (partial EuroBERT) support (#19826 )	2026-02-26 12:14:09 +01:00
llama-adapter.cpp	llama : re-enable manual LoRA adapter free (#19983 )	2026-03-18 12:03:26 +02:00
llama-adapter.h	llama : re-enable manual LoRA adapter free (#19983 )	2026-03-18 12:03:26 +02:00
llama-arch.cpp	model: mistral small 4 support (#20649 )	2026-03-17 00:31:14 +01:00
llama-arch.h	model: mistral small 4 support (#20649 )	2026-03-17 00:31:14 +01:00
llama-batch.cpp	kv-cache : fix M-RoPE checkpoints (#20132 )	2026-03-06 08:46:51 +02:00
llama-batch.h	batch : fix sequence id ownership (#17915 )	2025-12-11 14:29:47 +02:00
llama-chat.cpp	docs : Minor cleanups (#19252 )	2026-02-02 08:38:55 +02:00
llama-chat.h	model : add EXAONE MoE (#18543 )	2026-01-13 23:28:38 +01:00
llama-context.cpp	context : use n_embd_out for pooled embedding extraction (#20840 )	2026-03-21 19:35:00 +02:00
llama-context.h	graph : fix KQ mask, lora, cvec reuse checks (#19644 )	2026-02-16 09:21:11 +02:00
llama-cparams.cpp	cparams : rename LLAMA_MAX_PARALLEL_SEQUENCES to LLAMA_MAX_SEQ (#14188 )	2025-06-15 10:08:58 +03:00
llama-cparams.h	llama : enable chunked fused GDN path (#20340 )	2026-03-11 22:46:40 +02:00
llama-ext.h	test-backend-ops: allow loading tests from file and parsing model operators into file (#19896 )	2026-03-12 13:26:00 +01:00
llama-grammar.cpp	common/grammar: fix grammar parsing issues to prevent stack overflow and hangs (#18604 )	2026-03-21 18:43:35 +01:00
llama-grammar.h	common/grammar : replace problematic backtracking regex `[\s\S]*` (#18342 )	2026-01-03 16:02:43 -06:00
llama-graph.cpp	graph : add optional scale parameter to build_lora_mm [no ci] (#20427 )	2026-03-12 00:22:49 +01:00
llama-graph.h	graph : add optional scale parameter to build_lora_mm [no ci] (#20427 )	2026-03-12 00:22:49 +01:00
llama-hparams.cpp	llama: dynamic head_dim and n_rot for SWA (#20301 )	2026-03-09 22:22:39 +01:00
llama-hparams.h	llama : add support for Nemotron 3 Super (#20411 )	2026-03-11 19:27:53 +01:00
llama-impl.cpp	impl : use 6 digits for tensor dims (#20094 )	2026-03-04 09:53:38 +01:00
llama-impl.h	llama : enable chunked fused GDN path (#20340 )	2026-03-11 22:46:40 +02:00
llama-io.cpp	llama : refactor llama_context, llama_kv_cache, llm_build_context (#12181 )	2025-03-13 12:35:44 +02:00
llama-io.h	llama : refactor llama_context, llama_kv_cache, llm_build_context (#12181 )	2025-03-13 12:35:44 +02:00
llama-kv-cache-iswa.cpp	model : support Step3.5-Flash (#19283 )	2026-02-06 21:06:14 +01:00
llama-kv-cache-iswa.h	llama: print memory breakdown on exit (#15860 )	2025-09-24 16:53:48 +02:00
llama-kv-cache.cpp	kv-cache : fix reading llama_kv_cell_ext during state read (#20273 )	2026-03-15 09:11:19 +02:00
llama-kv-cache.h	llama: dynamic head_dim and n_rot for SWA (#20301 )	2026-03-09 22:22:39 +01:00
llama-kv-cells.h	llama: store mrope data in KV cell (#16825 )	2025-10-29 18:09:18 +01:00
llama-memory-hybrid-iswa.cpp	memory : add llama_memory_hybrid_iswa (#18601 )	2026-01-21 14:30:23 +02:00
llama-memory-hybrid-iswa.h	memory : add llama_memory_hybrid_iswa (#18601 )	2026-01-21 14:30:23 +02:00
llama-memory-hybrid.cpp	graph : reuse SSM graphs (#16490 )	2025-12-16 09:36:21 +02:00
llama-memory-hybrid.h	llama: print memory breakdown on exit (#15860 )	2025-09-24 16:53:48 +02:00
llama-memory-recurrent.cpp	server : support multi-modal context checkpoints (#19849 )	2026-02-25 15:14:27 +02:00
llama-memory-recurrent.h	llama: consistent ctx <-> buf order for KV cache (#16746 )	2025-10-28 11:23:54 +01:00
llama-memory.cpp	memory : correctly handle failure in apply() (#14438 )	2025-06-30 18:03:03 +03:00
llama-memory.h	llama: print memory breakdown on exit (#15860 )	2025-09-24 16:53:48 +02:00
llama-mmap.cpp	mmap: Fix Windows handle lifetime (#19598 )	2026-02-14 10:05:12 +02:00
llama-mmap.h	llama : add `use_direct_io` flag for model loading (#18166 )	2026-01-08 08:35:30 +02:00
llama-model-loader.cpp	ggml : add NVFP4 quantization type support (#19769 )	2026-03-11 21:02:54 +01:00
llama-model-loader.h	llama: end-to-end tests (#19802 )	2026-03-08 12:30:21 +01:00
llama-model-saver.cpp	llama: dynamic head_dim and n_rot for SWA (#20301 )	2026-03-09 22:22:39 +01:00
llama-model-saver.h	llama: end-to-end tests (#19802 )	2026-03-08 12:30:21 +01:00
llama-model.cpp	model : fix Granite Hybrid type check for 7B.A1B (#20795 )	2026-03-20 15:16:09 +01:00
llama-model.h	model : wire up Nemotron-H tensors for NVFP4 support (#20561 )	2026-03-16 09:19:16 +01:00
llama-quant.cpp	llama-quant : correct `n_attention_wv` usage (#20357 )	2026-03-10 21:43:29 +02:00
llama-quant.h	llama : refactor `src/llama.cpp` (#10902 )	2025-01-03 10:18:53 +02:00
llama-sampler.cpp	llama : rename llama-sampling to llama-sampler (#19363 )	2026-02-06 07:26:54 +01:00
llama-sampler.h	llama : rename llama-sampling to llama-sampler (#19363 )	2026-02-06 07:26:54 +01:00
llama-vocab.cpp	vocab : assert array size of scores and toktypes (#20737 )	2026-03-19 08:34:04 +01:00
llama-vocab.h	model : add JAIS-2 architecture support (#19488 )	2026-02-19 13:30:17 +01:00
llama.cpp	llama: end-to-end tests (#19802 )	2026-03-08 12:30:21 +01:00
unicode-data.cpp	server : better security control for public deployments (#9776 )	2024-10-08 13:27:04 +02:00
unicode-data.h	llama : reduce compile time and binary size (#9712 )	2024-10-02 15:49:55 +02:00
unicode.cpp	chore : correct typos [no ci] (#20041 )	2026-03-05 08:50:21 +01:00
unicode.h	devops: add s390x & ppc64le CI (#15925 )	2025-09-27 02:03:33 +08:00