llama.cpp

Commit Graph

Author	SHA1	Message	Date
Imad Saddik	22731da153	chore: update webui build output	2026-01-03 09:23:16 +01:00
Imad Saddik	eb997f61f9	Add autoChatWidth and customChatWidth to syncable parameters	2026-01-03 09:21:55 +01:00
Imad Saddik	21b35be366	chore: update webui build output	2026-01-03 09:12:29 +01:00
Imad Saddik	c9b34bc00d	Replace getChatWidth utility with chatWidthClasses in chat components	2026-01-03 09:10:44 +01:00
Imad Saddik	b6536f6589	Format code	2026-01-03 08:41:30 +01:00
Imad Saddik	36f334f4af	Put the constant into constants/chat-width.ts	2026-01-03 08:40:21 +01:00
Imad Saddik	00e6cafda6	Rename component to ChatSettingsComboboxCustomWidth	2026-01-03 08:37:09 +01:00
Imad Saddik	0143112cf0	chore: update webui build output	2026-01-03 08:35:42 +01:00
Imad Saddik	d7528b41fa	Merge remote-tracking branch 'upstream/master' into feat/change_chat_screen_width	2026-01-03 08:34:29 +01:00
tt	ced765be44	model: support youtu-vl model (#18479 ) * Support Youtu-VL Model * merge code * fix bug * revert qwen2 code & support rsplit in minja.hpp * update warm info * fix annotation * u * revert minja.hpp * fix * Do not write routed_scaling_factor to gguf when routed_scaling_factor is None * fix expert_weights_scale * LGTM after whitespace fixes * fix * fix * fix * layers to layer_index * enum fix --------- Co-authored-by: Xuan-Son Nguyen <son@huggingface.co> Co-authored-by: Sigbjørn Skjæret <sigbjorn.skjaeret@scala.com>	2026-01-01 19:25:54 +01:00
Anri Lombard	d5574c919c	webui: fix code copy stripping XML/HTML tags (#18518 ) * webui: fix code copy stripping XML/HTML tags * webui: update static build	2026-01-01 13:44:11 +01:00
Anri Lombard	33ded988ba	quantize: prevent input/output file collision (#18451 ) Check if input and output files are the same before quantizing to prevent file corruption when mmap reads from a file being written to. Fixes #12753	2025-12-31 23:29:03 +08:00
Henry147147	9b8329de7a	mtmd : Adding support for Nvidia Music Flamingo Model (#18470 ) * Inital commit, debugging q5_k_s quant * Made hf_to_gguf extend whisper to reduce code duplication * addressed convert_hf_to_gguf pull request issue --------- Co-authored-by: Henry D <henrydorsey147@gmail.com>	2025-12-31 12:13:23 +01:00
Jeff Bolz	f14f4e421b	server: fix files built redundantly (#18474 )	2025-12-30 13:11:13 +01:00
Xuan-Son Nguyen	51a48720b8	webui: fix prompt progress ETA calculation (#18468 ) * webui: fix prompt progress ETA calculation * handle case done === 0	2025-12-29 21:42:11 +01:00
Pascal	c9a3b40d65	Webui/prompt processing progress (#18300 ) * webui: display prompt preprocessing progress * webui: add percentage/ETA and exclude cached tokens from progress Address review feedback from ngxson * webui: add minutes and first chunk (0%) case * Update tools/server/webui/src/lib/components/app/chat/ChatMessages/ChatMessageAssistant.svelte Co-authored-by: Aleksander Grygier <aleksander.grygier@gmail.com> * Update tools/server/webui/src/lib/components/app/chat/ChatMessages/ChatMessageAssistant.svelte Co-authored-by: Aleksander Grygier <aleksander.grygier@gmail.com> * webui: address review feedback from allozaur * chore: update webui build output * webui: address review feedback from allozaur * nit * chore: update webui build output * feat: Enhance chat processing state * feat: Improve chat processing statistics UI * chore: update webui build output * feat: Add live generation statistics to processing state hook * feat: Persist prompt processing stats in hook for better UX * refactor: Enhance ChatMessageStatistics for live stream display * feat: Implement enhanced live chat statistics into assistant message * chore: update webui build output * fix: Proper tab for each stage of prompt processing/generation * chore: update webui build output * fix: Improved ETA calculation & display logic * chore: update webui build output * feat: Simplify logic & remove ETA from prompt progress * chore: update webui build output --------- Co-authored-by: Aleksander Grygier <aleksander.grygier@gmail.com>	2025-12-29 19:32:21 +01:00
wbtek	5b1248c9af	server : Cmdline arg -to changes http read timeout from current 600sec default (#18279 ) * Prevent crash if TTFT >300sec, boosted to 90 days * server : allow configurable HTTP timeouts for child models * server : pass needed timeouts from params only --------- Co-authored-by: Greg Slocum <fromgit@wbtek.slocum.net>	2025-12-29 17:12:48 +01:00
Georgi Gerganov	2a85f720b8	server : handle closed connection for tasks (#18459 )	2025-12-29 15:34:41 +02:00
o7si	daa242dfc8	common: fix return value check for setpriority (#18412 ) * common: fix return value check for setpriority * tools: add logging for process priority setting	2025-12-29 11:07:49 +02:00
Xuan-Son Nguyen	cffa5c46ea	mtmd: clarify that we no longer accept AI-generated PRs (#18406 )	2025-12-28 09:57:04 +01:00
Johannes Gäßler	a52dc60ba3	llama_fit_params: return enum for fail vs. error (#18374 )	2025-12-27 09:59:19 +01:00
o7si	4893cc07bb	server : fix crash when seq_rm fails for hybrid/recurrent models (#18391 ) * server : fix crash when seq_rm fails for hybrid/recurrent models * server : add allow_processing param to clear_slot	2025-12-26 16:35:29 +01:00
Xuan-Son Nguyen	f5acfb2ffa	server: (router) add stop-timeout option (#18350 ) * server: (router) add stop-timeout option * also allow stop while loading * add docs * unload_lru: also wait for unload to complete	2025-12-24 23:47:49 +01:00
Aadeshveer Singh	c184284230	fit-params : fix race condition in fit-params output (#18276 )	2025-12-24 15:57:38 +01:00
Xuan-Son Nguyen	5ee4e43f26	server: return_progress to also report 0% processing state (#18305 )	2025-12-23 21:49:05 +01:00
Pascal	5b6c9bc0f3	webui: apply webui_settings on first load (#18223 ) * webui: apply webui_settings on first load The webui_settings from /props were not applied on initial load when default_generation_settings.params was null Now syncs whenever serverProps is available, regardless of params, works for both single-model and router modes * chore: update webui build output	2025-12-23 15:48:03 +01:00
Xuan-Son Nguyen	849d021104	server: fix crash with model not having BOS/EOS (#18321 )	2025-12-23 14:39:36 +01:00
Xuan-Son Nguyen	179fd82a72	gen-docs: automatically update markdown file (#18294 ) * gen-docs: automatically update markdown file * also strip whitespace * do not add extra newline * update TOC	2025-12-22 19:30:19 +01:00
Xuan-Son Nguyen	6ce863c803	server: prevent data race from HTTP threads (#18263 ) * server: prevent data race from HTTP threads * fix params * fix default_generation_settings * nits: make handle_completions_impl looks less strange * stricter const * fix GGML_ASSERT(idx < states.size()) * move index to be managed by server_response_reader * http: make sure req & res lifecycle are tied together * fix compile * fix index handling buggy * fix data race for lora endpoint * nits: fix shadow variable * nits: revert redundant changes * nits: correct naming for json_webui_settings	2025-12-22 14:23:34 +01:00
Xuan-Son Nguyen	3997c78e33	server: fix data race in to_json_anthropic (#18283 )	2025-12-22 13:21:43 +01:00
Xuan-Son Nguyen	86af848153	server: (docs) remove mention about extra_args (#18262 )	2025-12-22 12:22:01 +01:00
Johannes Gäßler	147a521636	tool/ex/tests: consistently free ctx, then model (#18168 )	2025-12-22 11:00:37 +01:00
Xuan-Son Nguyen	ddcb75dd8a	server: add auto-sleep after N seconds of idle (#18228 ) * implement sleeping at queue level * implement server-context suspend * add test * add docs * optimization: add fast path * make sure to free llama_init * nits * fix use-after-free * allow /models to be accessed during sleeping, fix use-after-free * don't allow accessing /models during sleep, it is not thread-safe * fix data race on accessing props and model_meta * small clean up * trailing whitespace * rm outdated comments	2025-12-21 02:24:42 +01:00
Imad Saddik	642a4d68b8	chore: update webui build output	2025-12-20 15:44:02 +01:00
Imad Saddik	9b87eaf898	Updated the disabled text	2025-12-20 15:42:21 +01:00
Imad Saddik	fac6ef71a8	chore: update webui build output	2025-12-20 15:27:35 +01:00
Imad Saddik	37be01ac4c	Applied formatting	2025-12-20 15:26:23 +01:00
Imad Saddik	689b7a5bd6	Removed console log and refactored the code	2025-12-20 15:25:41 +01:00
Imad Saddik	6db71973ab	Applied formatting	2025-12-20 15:18:41 +01:00
Imad Saddik	b519235e0a	Added the ability to set custom pixel values in the combobox	2025-12-20 15:17:52 +01:00
Imad Saddik	3be006dcaf	Display the custom chat width combobox in the settings with just presets	2025-12-20 14:29:15 +01:00
Imad Saddik	14929a77b0	Added the command and popover components	2025-12-20 11:10:06 +01:00
Oleksandr Kuvshynov	408616adbd	server : [easy] fix per round speculative decode logging (#18211 ) Currently we always log 0, as we clear slot.drafted before. To reproduce: Run llama-server with devstral-2 as main model and devstral-2-small as md, and verbose logging: ``` % ./build/bin/llama-server -v \ -m ~/llms/Devstral-2-123B-Instruct-2512-UD-Q6_K_XL-00001-of-00003.gguf \ -md ~/llms/Devstral-Small-2-24B-Instruct-2512-UD-Q2_K_XL.gguf \ -c 8192 2> /tmp/llama.cpp.debug Check the log: slot update_slots: id 3 \| task 0 \| accepted 11/0 draft tokens, new n_tokens = 741 slot update_slots: id 3 \| task 0 \| accepted 4/0 draft tokens, new n_tokens = 746 slot update_slots: id 3 \| task 0 \| accepted 16/0 draft tokens, new n_tokens = 763 slot update_slots: id 3 \| task 0 \| accepted 11/0 draft tokens, new n_tokens = 775 slot update_slots: id 3 \| task 0 \| accepted 2/0 draft tokens, new n_tokens = 778 slot update_slots: id 3 \| task 0 \| accepted 4/0 draft tokens, new n_tokens = 783 slot update_slots: id 3 \| task 0 \| accepted 8/0 draft tokens, new n_tokens = 792 slot update_slots: id 3 \| task 0 \| accepted 2/0 draft tokens, new n_tokens = 795 slot update_slots: id 3 \| task 0 \| accepted 1/0 draft tokens, new n_tokens = 797 slot update_slots: id 3 \| task 0 \| accepted 1/0 draft tokens, new n_tokens = 799 slot update_slots: id 3 \| task 0 \| accepted 0/0 draft tokens, new n_tokens = 800 slot update_slots: id 3 \| task 0 \| accepted 2/0 draft tokens, new n_tokens = 803 slot update_slots: id 3 \| task 0 \| accepted 1/0 draft tokens, new n_tokens = 805 slot update_slots: id 3 \| task 0 \| accepted 6/0 draft tokens, new n_tokens = 812 slot update_slots: id 3 \| task 0 \| accepted 3/0 draft tokens, new n_tokens = 816 ``` After the fix, get correct per round logging: ``` slot update_slots: id 3 \| task 0 \| accepted 7/8 draft tokens, new n_tokens = 654 slot update_slots: id 3 \| task 0 \| accepted 1/2 draft tokens, new n_tokens = 656 slot update_slots: id 3 \| task 0 \| accepted 2/16 draft tokens, new n_tokens = 659 slot update_slots: id 3 \| task 0 \| accepted 1/16 draft tokens, new n_tokens = 661 slot update_slots: id 3 \| task 0 \| accepted 2/16 draft tokens, new n_tokens = 664 slot update_slots: id 3 \| task 0 \| accepted 16/16 draft tokens, new n_tokens = 681 slot update_slots: id 3 \| task 0 \| accepted 16/16 draft tokens, new n_tokens = 698 slot update_slots: id 3 \| task 0 \| accepted 3/4 draft tokens, new n_tokens = 702 slot update_slots: id 3 \| task 0 \| accepted 5/12 draft tokens, new n_tokens = 708 slot update_slots: id 3 \| task 0 \| accepted 16/16 draft tokens, new n_tokens = 725 slot update_slots: id 3 \| task 0 \| accepted 1/1 draft tokens, new n_tokens = 727 slot update_slots: id 3 \| task 0 \| accepted 8/16 draft tokens, new n_tokens = 736 ```	2025-12-20 10:57:40 +01:00
Imad Saddik	416bb35130	Used the new getChatWidth function in ChatProcessingInfo	2025-12-20 09:28:55 +01:00
Imad Saddik	62614a5faa	Used the new getChatWidth function in all chat messages components that need it	2025-12-20 09:27:49 +01:00
Xuan-Son Nguyen	9e39a1e6a9	server: support load model on startup, support preset-only options (#18206 ) * server: support autoload model, support preset-only options * add docs * load-on-startup * fix * Update common/arg.cpp Co-authored-by: Pascal <admin@serveurperso.com> --------- Co-authored-by: Pascal <admin@serveurperso.com>	2025-12-20 09:25:27 +01:00
Imad Saddik	dccfcc02eb	Used the new getChatWidth function in ChatForm	2025-12-20 09:09:45 +01:00
Imad Saddik	33d8d0f461	Performed formatting	2025-12-20 09:07:54 +01:00
Imad Saddik	61d99bbd88	Used the new chat width logic in ChatScreen and ChatWarning	2025-12-20 09:07:04 +01:00
Imad Saddik	b6dbbcc1fb	Moved and renamed the width-classes.ts file	2025-12-20 09:04:13 +01:00

1 2 3 4 5 ...

493 Commits