llama.cpp

Commit Graph

Author	SHA1	Message	Date
Francis Couture-Harpin	2ef41855cf	convert : for FP8, use scale type to decide auto type	2025-11-06 22:55:53 -05:00
Francis Couture-Harpin	f88a4b9398	gguf-py : handle cross-filesystem file range copies	2025-11-06 22:55:53 -05:00
Francis Couture-Harpin	4be1a5d44b	convert : better logging of partially reflinkable tensors	2025-11-06 22:55:53 -05:00
Francis Couture-Harpin	6ffa46d8f4	gguf-py : allow previewing reflinked size on non-Linux platforms	2025-11-06 22:55:53 -05:00
Francis Couture-Harpin	3126b5ee4e	convert : remove unused field ModelTensorInfo.src_qtype	2025-11-06 22:55:53 -05:00
Francis Couture-Harpin	e097d98a22	convert : more robust default ftype detection	2025-11-06 22:55:53 -05:00
Francis Couture-Harpin	5712aa895f	gguf-py : improve reflink size logging * gguf-py : move reflinking functions to lazy	2025-11-06 22:55:53 -05:00
Francis Couture-Harpin	d3fcb0e90e	convert : allow sharding reflinked models	2025-11-06 22:55:53 -05:00
Francis Couture-Harpin	614b95a88d	convert : use F32 operations on Mamba A_log This matches the previous behavior for BF16 tensors.	2025-11-06 22:55:53 -05:00
Francis Couture-Harpin	c3738cfcef	convert : detect filesystem block size for reflinks * convert : use direct copies when possible Using os.copy_file_range where available, and falling back to shutil.copyfileobj otherwise. * gguf : handle misaligned offset more cleanly	2025-11-06 22:55:53 -05:00
Francis Couture-Harpin	791bd97b3c	gguf-py : fix flake8 lint	2025-11-06 22:55:53 -05:00
Francis Couture-Harpin	d921057027	convert : fix reflinks for stacked MoE tensors	2025-11-06 22:55:53 -05:00
Francis Couture-Harpin	562aa42c12	convert : use reflinks for faster conversion	2025-11-06 22:55:52 -05:00
Francis Couture-Harpin	e996f3aef8	convert : fix no-lazy dtypes from direct safetensors	2025-11-06 22:33:09 -05:00
Francis Couture-Harpin	e7b7ed8ab1	gguf-py : order safetensors tensors by name Applies to both local and remote safetensors custom parsing. This matches the behavior of the official safetensors implementation. * convert : rename from_safetensors_meta to from_local_tensor For consistency with from_remote_tensor	2025-11-06 22:33:09 -05:00
Francis Couture-Harpin	c4b630f25d	convert : parse safetensors directly	2025-11-06 22:33:09 -05:00
xctan	7f09a680af	ggml-cpu : optimize RVV q2_k and q3_k kernels (#16887 )	2025-11-06 18:12:45 +02:00
Johannes Gäßler	aa374175c3	CUDA: fix crash on uneven context without FA (#16988 )	2025-11-06 14:05:47 +01:00
Georgi Gerganov	5b180c3d60	metal : initial Metal4 tensor API support (#16634 ) * metal : rework mat-mat multiplication * metal : initial Metal4 support * cont * metal : detect tensor support * cont : better ifdefs * metal : support tensors in mul_mm_id * metal : add env for disabling tensor API * tests : restore * metal : remove unused constants * metal : fix check for bfloat tensor support * cont : handle API incompatibilities * cont : handle even more incompatibilities * metal : use tensor API only on M5 and later	2025-11-06 14:45:10 +02:00
Georgi Gerganov	b7f9010d24	server : disable checkpoints with mtmd (#17045 )	2025-11-06 12:09:29 +02:00
Xuan-Son Nguyen	4882f0ff78	clip: implement minicpm-v sinusoidal embd using GGML (#17036 ) * clip: implement minicpm-v sinusoidal embd using GGML * fix repeat op	2025-11-06 11:02:54 +01:00
YehuditE	9d7c518d64	sycl: add CONCAT operator support (#16047 ) * sycl: add CONCAT operator support * cleanup: remove stray lines added by mistake * fix: code format issues in concat.cpp and tests/test-backend-ops.cpp * chore: fix editorconfig violations * cleanup: drop unnecessary i16 type support * docs: update sycl-csv and regenerate ops.md * update docs/ops.md * fix: adapt to upstream master changes after rebase * fix: remove empty files * fix: drop whitespace --------- Co-authored-by: Sigbjørn Skjæret <sigbjorn.skjaeret@scala.com>	2025-11-06 11:02:33 +01:00
Johannes Gäßler	22c8c3c6ad	docs: explain CUDA 11 compilation [no ci] (#16824 )	2025-11-06 08:14:35 +01:00
l3utterfly	6db3d1ffe6	ggml-hexagon: graceful fallback for older socs where rpcmem_alloc2 and FASTRPC_GET_URI is unsupported (#16987 ) * support older socs where FASTRPC_GET_URI is unsupported * added graceful fallback when FASTRPC_GET_URI call fails * use weak symbols instead of loading libcdsprpc.so dynamically * Add weak pragma for rpcmem_alloc2 * Remove weak declaration for rpcmem_alloc2 in ggml-hexagon.cpp Removed weak declaration for rpcmem_alloc2. * Enforce ndev to 1 for archs below v75 Force ndev to 1 for SoCs architectures lower than v75.	2025-11-05 21:46:38 -08:00
bssrdf	230d1169e5	improve CUDA cpy memory bandwidth when copying transposed tensor (#16841 ) * WIP * added a cpy kernel specific to transposed tensor which uses smem to avoid uncoalesced access; test cases also added shwoing improved memory bandwidth * added BF16 support * more strict check to make sure src0 is a transpose * reformulated to handle more complicated transpose cases * bring back 2D transpose for higher performance * allow build on windows * tranpose copy more shapes * minor tweak * final clean up * restore some test cases * keep only the kernel for true tranposed case; updated with review suggestions * make CI happy * remove headers not needed * reduced bank conflicts for fp16 and bf16 * add missing const* * now bank conflicts free * use padding instead of swizzling --------- Co-authored-by: bssrdf <bssrdf@gmail.com>	2025-11-05 21:55:04 +01:00
Jeff Bolz	a44d77126c	vulkan: Fix GGML_VULKAN_CHECK_RESULTS to better handle fusion (#16919 )	2025-11-05 19:51:03 +01:00
Gabe Goodhart	5886f4f545	examples(gguf): GGUF example outputs (#17025 ) * feat(llama-gguf): Print out the tensor type in llama-gguf r Branch: Mamba2Perf Signed-off-by: Gabe Goodhart <ghart@us.ibm.com> * feat(off-topic): print the number of elements in tensors with llama-gguf Branch: Mamba2SSD Signed-off-by: Gabe Goodhart <ghart@us.ibm.com> * style: valign Branch: GGUFToolOutputs Signed-off-by: Gabe Goodhart <ghart@us.ibm.com> * Update examples/gguf/gguf.cpp --------- Signed-off-by: Gabe Goodhart <ghart@us.ibm.com> Co-authored-by: Georgi Gerganov <ggerganov@gmail.com>	2025-11-05 19:58:16 +02:00
Xuan-Son Nguyen	92bb84f775	mtmd: allow QwenVL to process larger image by default (#17020 )	2025-11-05 14:26:49 +01:00
Georgi Gerganov	13b339bcd9	server : do not default to multiple slots with speculative decoding (#17017 ) * server : do not default to multiple slots with speculative decoding * cont : fix	2025-11-05 14:32:55 +02:00
Xuan-Son Nguyen	2f0c2db43e	mtmd: improve struct initialization (#16981 )	2025-11-05 11:26:37 +01:00
손희준	fd2f84f468	docs: Clarify the endpoint that webui uses (#17001 )	2025-11-05 11:20:28 +01:00
Li Pengzhan	9f052478c2	model : add openPangu-Embedded (#16941 ) * Model: add openPangu-Embedded * fixed according to reviewer's comments * fixed the chat template check condition * Apply suggestions from code review change the chat-template check condition and some formatting issue Co-authored-by: Sigbjørn Skjæret <sigbjorn.skjaeret@scala.com> * whitespace cleanup --------- Co-authored-by: Sigbjørn Skjæret <sigbjorn.skjaeret@scala.com>	2025-11-05 10:28:58 +01:00
Reese Levine	03ea04175d	ggml webgpu: minor set rows optimization (#16810 ) * Add buffer label and enable dawn-specific toggles to turn off some checks * Minor set_rows optimization (#4) * updated optimization, fixed errors * non vectorized version now dispatches one thread per element * Simplify * Change logic for set_rows pipelines --------- Co-authored-by: Neha Abbas <nehaabbas@macbookpro.lan> Co-authored-by: Neha Abbas <nehaabbas@ReeseLevines-MacBook-Pro.local> Co-authored-by: Reese Levine <reeselevine1@gmail.com> * Comment on dawn toggles * Remove some comments * Implement overlap binary operators * Revert "Implement overlap binary operators" This reverts commit `ed710b36f5`. * Disable support for non-contiguous binary_op tensors and leave note for future support --------- Co-authored-by: neha-ha <137219201+neha-ha@users.noreply.github.com> Co-authored-by: Neha Abbas <nehaabbas@macbookpro.lan> Co-authored-by: Neha Abbas <nehaabbas@ReeseLevines-MacBook-Pro.local>	2025-11-05 10:27:42 +01:00
Georgi Gerganov	cdabeb2c27	sync : ggml	2025-11-05 10:41:51 +02:00
Georgi Gerganov	852ce5180a	ggml : fix conv2d_dw SVE path (ggml/1380) * Fix test-conv2d-dw failure on ARM SVE by using runtime vector length The ggml_compute_forward_conv_2d_dw_cwhn function was using a hardcoded GGML_F32_EPR (8) for SIMD vectorization, but on ARM SVE the actual vector length varies by hardware. This caused incorrect computation when processing CWHN layout tensors on ARM machines. Fix by using svcntw() to get the runtime SVE vector length instead of the compile-time constant. Co-authored-by: ggerganov <1991296+ggerganov@users.noreply.github.com> * ci : reduce sam score threshold * ci : update bbox checks for sam test --------- Co-authored-by: copilot-swe-agent[bot] <198982749+Copilot@users.noreply.github.com> Co-authored-by: ggerganov <1991296+ggerganov@users.noreply.github.com>	2025-11-05 10:41:51 +02:00
mnehete32	9aa63374f2	CUDA: update ops.md (#17005 )	2025-11-05 11:01:15 +08:00
lhez	5e90233bdb	opencl: update doc (#17011 ) * opencl: update docs * opencl: update docs * opencl: fix link * opencl: update doc	2025-11-04 16:02:36 -08:00
nullname	a5c07dcd7b	refactor: replace sprintf with snprintf for safer string handling in dump functions (#16913 )	2025-11-04 12:25:39 -08:00
Jeff Bolz	ad51c0a720	vulkan: remove the need for the dryrun (#16826 ) * vulkan: remove the need for the dryrun Allocate pipelines and descriptor sets when requested. Reallocate the prealloc buffers when needed, and flush any pending work before reallocating. For rms_partials and total_mul_mat_bytes, use the sizes computed the last time the graph was executed. * remove dryrun parameters	2025-11-04 13:28:17 -06:00
Georgi Gerganov	66d8eccd42	server : do context shift only while generating (#17000 )	2025-11-04 19:21:36 +02:00
Georgi Gerganov	afd353246d	readme : update hot topics (#17002 )	2025-11-04 17:21:31 +02:00
Acly	cc98f8d349	ggml-cpu : bicubic interpolation (#16891 )	2025-11-04 13:12:20 +01:00
Sigbjørn Skjæret	d945834366	ci : apply model label to models (#16994 )	2025-11-04 12:29:39 +01:00
Sigbjørn Skjæret	b164259bba	chore : fix models indent after refactor (#16992 )	2025-11-04 12:29:15 +01:00
Noah	1f5accb8d0	Fix garbled output with REPACK at high thread counts (#16956 ) * Fix garbled output with REPACK at high thread counts Fixed a race condition in the REPACK matrix multiplication code that caused garbled output when using 26+ threads (model-dependent threshold). The issue occurred because with high thread counts, the code forced chunk count to equal thread count, creating many small chunks. After aligning these chunks to NB_COLS boundaries, adjacent chunks could overlap, causing data corruption and race conditions. The fix enforces minimum chunk sizes based on NB_COLS and caps maximum chunk count to prevent creating too many tiny chunks, ensuring proper alignment without overlaps. * Update ggml/src/ggml-cpu/repack.cpp Co-authored-by: Georgi Gerganov <ggerganov@gmail.com> * Update ggml/src/ggml-cpu/repack.cpp Co-authored-by: Georgi Gerganov <ggerganov@gmail.com> --------- Co-authored-by: Georgi Gerganov <ggerganov@gmail.com>	2025-11-03 21:04:59 -08:00
Aman Gupta	2759ccdb4a	CUDA: avoid mul + bias fusion when doing fusion (#16935 )	2025-11-04 10:53:48 +08:00
lhez	c5023daf60	opencl: support imrope (#16914 ) * opencl: support imrope * opencl: fix whitespace	2025-11-03 11:47:57 -08:00
Aleksander Grygier	e7da30b584	fix: Viewing multiple PDF attachments (#16974 )	2025-11-03 18:53:26 +01:00
Daniel Bevenius	ed8aa63320	model-conversion : pass config to from_pretrained (#16963 ) This commit modifies the script `run-org-model.py` to ensure that the model configuration is explicitly passed to the `from_pretrained` method when loading the model. It also removes a duplicate configuration loading which was a mistake. The motivation for this change is that enables the config object to be modified and then passed to the model loading function, which can be useful when testing new models.	2025-11-03 18:01:59 +01:00
Georgi Gerganov	48bd26501b	server : add props.model_alias (#16943 ) * server : add props.model_alias * webui : npm run format	2025-11-03 14:38:23 +01:00

1 2 3 4 5 ...

6986 Commits All Branches Search

6986 Commits

All Branches