llama.cpp

History

nullname 379bdeb18c feat: perf opt gemv (#54 ) * add GEMV implementation for matrix multiplication in hexagon * refactor: optimize GEMV implementation for matrix multiplication in hexagon * wip * refactor: enhance caching mechanism in GEMV implementation for matrix multiplication * wip * refactor: streamline caching logic in GEMV implementation for matrix multiplication * wip * wip * fix broadcase in flash_attn * format * refactor: optimize memory fetching in matrix multiplication implementations * wip * fix aligned gemv * rename * refactor: remove unused memory cache functions and initialize VTCM cache * wip * feat: add vector math functions for IEEE float and half float operations * feat: add vec_silu_f32 and vec_silu_f16 functions for SiLU activation * feat: implement GLU operation support in tensor processing * feat: add GLU operation support and related enhancements in tensor processing * wip * wip * wip * feat: add qhmath_hvx_div_vf functions for f32 vector operations * feat: add qhmath_hvx_div_vhf functions for f16 vector operations * fix: reorder parameters in vector operation functions for consistency * wip * feat: enhance vector operations with parameterized transformations and improved GLU implementations * wip * fix: increase default stack size and correct thread parameter indexing in thread pool * fix f16 div * fix f32 div * fix: update GLU vector operations to use explicit denominator calculation * wip * wip * Refactor cacheability check for matrix multiplication to handle multiple source tensors * Revert "fix: increase default stack size and correct thread parameter indexing in thread pool" This reverts commit 40e3f0974dbb04051aa30b397a9a171c6dd32678. * wip * fix comments * replace copy with memcpy		2025-08-08 20:40:26 +08:00
..
ggml-blas	cmake : Fix broken CMake error messages (ggml/1252)	2025-06-01 13:43:57 +03:00
ggml-cann	docker : add cann build pipline (#14591 )	2025-08-01 10:02:34 +08:00
ggml-cpu	ggml : Q2k interleaving implementation - x86/x64 SIMD (#14373 )	2025-08-01 09:20:33 +03:00
ggml-cuda	CUDA: fix MMQ nwarps for AMD with warp_size==32 (#15014 )	2025-08-01 20:47:32 +02:00
ggml-hip	HIP: add GGML_HIP_MMQ_MFMA option to allow disableing the MFMA path. (#14930 )	2025-07-29 17:44:30 +02:00
ggml-metal	metal: SSM_SCAN performance (#14743 )	2025-07-25 10:47:39 -06:00
ggml-musa	musa: upgrade musa sdk to rc4.2.0 (#14498 )	2025-07-24 20:05:37 +01:00
ggml-opencl	opencl: add f16 for `add`, `sub`, `mul`, `div` (#14984 )	2025-08-01 13:15:44 +02:00
ggml-qnn	feat: perf opt gemv (#54 )	2025-08-08 20:40:26 +08:00
ggml-rpc	rpc : check for null buffers in get/set/copy tensor endpoints (#14868 )	2025-07-25 12:17:02 +02:00
ggml-sycl	SYCL: Add set_rows support for quantized types (#14883 )	2025-07-28 20:32:15 +05:30
ggml-vulkan	Vulkan: Fix minor debug mode issues (#14899 )	2025-07-31 17:46:54 +02:00
ggml-webgpu	ggml: Add initial WebGPU backend (#14521 )	2025-07-16 18:18:51 +03:00
CMakeLists.txt	Merge branch 'master' into dev-refactoring	2025-07-18 23:43:20 +08:00
ggml-alloc.c	metal : fuse add, mul + add tests (#14596 )	2025-07-18 20:37:26 +03:00
ggml-backend-impl.h	ggml : upgrade init_tensor API to return a ggml_status (#11854 )	2025-02-28 14:41:47 +01:00
ggml-backend-reg.cpp	Merge branch 'master' into dev-refactoring	2025-07-18 23:43:20 +08:00
ggml-backend.cpp	sched : fix multiple evaluations of the same graph with pipeline parallelism (#14855 )	2025-07-25 11:07:26 +03:00
ggml-common.h	ggml-cpu : split arch-specific implementations (#13892 )	2025-06-09 16:47:13 +02:00
ggml-impl.h	metal : fuse add, mul + add tests (#14596 )	2025-07-18 20:37:26 +03:00
ggml-opt.cpp	mnist: fix segmentation fault (ggml/1227)	2025-05-19 13:29:56 +03:00
ggml-quants.c	ggml-quants : rename best_mad to best_error (ggml/1283)	2025-07-01 11:06:39 +03:00
ggml-quants.h	ggml : build backends as libraries (#10256 )	2024-11-14 18:04:35 +01:00
ggml-threading.cpp	ggml : build backends as libraries (#10256 )	2024-11-14 18:04:35 +01:00
ggml-threading.h	remove CMAKE_WINDOWS_EXPORT_ALL_SYMBOLS (#10797 )	2024-12-12 19:02:49 +01:00
ggml.c	ggml : remove invalid portPos specifiers from dot files (#14838 )	2025-07-25 14:29:57 +03:00
ggml.cpp	ggml : Print backtrace on uncaught C++ exceptions (ggml/1232)	2025-06-01 13:43:57 +03:00
gguf.cpp	ggml : prevent integer overflow in gguf tensor size calculation (#14595 )	2025-07-09 14:33:53 +02:00