llama.cpp/ggml/src
Jeff Bolz 4cb208c93c
vulkan: coopmat2 mul_mat optimizations (#14934)
- Increase tile size for k-quants, to match non-k-quants
- Choose more carefully between large and medium tiles, considering how it
  interacts with split_k
- Allow larger/non-power of two split_k, and make the splits a multiple of 256
- Use split_k==3 to when >1/2 and <=2/3 of the SMs would hae been used
2025-08-02 11:21:37 +02:00
..
ggml-blas cmake : Fix broken CMake error messages (ggml/1252) 2025-06-01 13:43:57 +03:00
ggml-cann docker : add cann build pipline (#14591) 2025-08-01 10:02:34 +08:00
ggml-cpu ggml : Q2k interleaving implementation - x86/x64 SIMD (#14373) 2025-08-01 09:20:33 +03:00
ggml-cuda CUDA: fix MMQ nwarps for AMD with warp_size==32 (#15014) 2025-08-01 20:47:32 +02:00
ggml-hip HIP: add GGML_HIP_MMQ_MFMA option to allow disableing the MFMA path. (#14930) 2025-07-29 17:44:30 +02:00
ggml-metal metal: SSM_SCAN performance (#14743) 2025-07-25 10:47:39 -06:00
ggml-musa musa: upgrade musa sdk to rc4.2.0 (#14498) 2025-07-24 20:05:37 +01:00
ggml-opencl opencl: add f16 for `add`, `sub`, `mul`, `div` (#14984) 2025-08-01 13:15:44 +02:00
ggml-rpc rpc : check for null buffers in get/set/copy tensor endpoints (#14868) 2025-07-25 12:17:02 +02:00
ggml-sycl SYCL: Add set_rows support for quantized types (#14883) 2025-07-28 20:32:15 +05:30
ggml-vulkan vulkan: coopmat2 mul_mat optimizations (#14934) 2025-08-02 11:21:37 +02:00
ggml-webgpu ggml: Add initial WebGPU backend (#14521) 2025-07-16 18:18:51 +03:00
CMakeLists.txt ggml: Add initial WebGPU backend (#14521) 2025-07-16 18:18:51 +03:00
ggml-alloc.c metal : fuse add, mul + add tests (#14596) 2025-07-18 20:37:26 +03:00
ggml-backend-impl.h ggml : upgrade init_tensor API to return a ggml_status (#11854) 2025-02-28 14:41:47 +01:00
ggml-backend-reg.cpp ggml: Add initial WebGPU backend (#14521) 2025-07-16 18:18:51 +03:00
ggml-backend.cpp sched : fix multiple evaluations of the same graph with pipeline parallelism (#14855) 2025-07-25 11:07:26 +03:00
ggml-common.h ggml-cpu : split arch-specific implementations (#13892) 2025-06-09 16:47:13 +02:00
ggml-impl.h metal : fuse add, mul + add tests (#14596) 2025-07-18 20:37:26 +03:00
ggml-opt.cpp mnist: fix segmentation fault (ggml/1227) 2025-05-19 13:29:56 +03:00
ggml-quants.c ggml-quants : rename best_mad to best_error (ggml/1283) 2025-07-01 11:06:39 +03:00
ggml-quants.h ggml : build backends as libraries (#10256) 2024-11-14 18:04:35 +01:00
ggml-threading.cpp ggml : build backends as libraries (#10256) 2024-11-14 18:04:35 +01:00
ggml-threading.h remove CMAKE_WINDOWS_EXPORT_ALL_SYMBOLS (#10797) 2024-12-12 19:02:49 +01:00
ggml.c ggml : remove invalid portPos specifiers from dot files (#14838) 2025-07-25 14:29:57 +03:00
ggml.cpp ggml : Print backtrace on uncaught C++ exceptions (ggml/1232) 2025-06-01 13:43:57 +03:00
gguf.cpp ggml : prevent integer overflow in gguf tensor size calculation (#14595) 2025-07-09 14:33:53 +02:00