llama.cpp

History

nullname 85dde8dc4a hexagon: optimize HMX matmul operations (#21071 ) * optimize hmx_mat_mul functions by calculating row and column tiles upfront * refactor core_dot_chunk_fp16 to use size_t for tile counts and improve readability * wip * set scale outside of loop * wip * refactor core_mma_chunk_fp16 and mat_mul_qk_0_d16a32 to use size_t for tile counts * wip * wip * refactor transfer_output_chunk_fp16_to_fp32 to use size_t for dimensions * refactor core_dot_chunk_fp16 to use size_t for tile row stride calculation * wip * refactor hmx_mat_mul functions to use hvx_vec_splat_f16 for column scales initialization * refactor hmx_mat_mul_permuted_w16a32_batched to streamline scale setting and locking * refactor core_dot_chunk_fp16 to improve tile stride calculations for output * refactor hmx_mat_mul functions to use Q6_V_vsplat_R for column scales initialization * fix compiling error * wip * optimize row and column tile indexing in core_mma_chunk_fp16 function * wip * Revert "wip" This reverts commit `cde679eff7`. * Add size limit check for HAP_mmap in htp_iface_mmap and drop_mmap functions * wip		2026-04-16 13:48:34 -07:00
..
htp	hexagon: optimize HMX matmul operations (#21071 )	2026-04-16 13:48:34 -07:00
CMakeLists.txt	ggml-hexagon: flash-attention and reduce-sum optimizations (#19141 )	2026-01-30 21:14:20 -08:00
ggml-hexagon.cpp	hexagon: improved Op queuing, buffer and cache management (#21705 )	2026-04-10 15:47:43 -07:00
htp-drv.cpp	chore : correct typos [no ci] (#20041 )	2026-03-05 08:50:21 +01:00
htp-drv.h	hexagon: enable offloading to Hexagon on Windows on Snapdragon (#19150 )	2026-01-29 12:33:21 -08:00
libdl.h	hexagon: enable offloading to Hexagon on Windows on Snapdragon (#19150 )	2026-01-29 12:33:21 -08:00
libggml-htp.inf	hexagon: enable offloading to Hexagon on Windows on Snapdragon (#19150 )	2026-01-29 12:33:21 -08:00
op-desc.h	ggml-hexagon: create generalized functions for cpu side op (#17500 )	2025-12-22 23:13:24 -08:00