llama.cpp

History

Jeff Bolz be47fb9285 vulkan: extend topk_moe to handle sigmoid w/exp_probs_b for nemotron (#18295 ) * vulkan: extend topk_moe to handle sigmoid w/exp_probs_b for nemotron Also handle GGML_OP_SCALE at the end (nemotron, deepseek2). Fewer pipeline variants and spec constants, just use push constants. In test_topk_moe, change exp_probs_b to be 1D, matching real networks. Update test-backend-ops and ggml-backend to allow verifying multiple outputs in a fusion test (topk_moe has two outputs). Previously only the final node was verified. * change test_topk_moe to allow results in arbitrary order * disable sigmoid fusion for moltenvk		2026-01-01 08:58:27 +01:00
..
cmake	cmake: fix ggml-shaders-gen compiler paths containing spaces (#12747 )	2025-04-04 10:12:40 -03:00
vulkan-shaders	vulkan: extend topk_moe to handle sigmoid w/exp_probs_b for nemotron (#18295 )	2026-01-01 08:58:27 +01:00
CMakeLists.txt	vulkan: Improve build time for MSVC (#16545 )	2025-10-14 14:51:36 +02:00
ggml-vulkan.cpp	vulkan: extend topk_moe to handle sigmoid w/exp_probs_b for nemotron (#18295 )	2026-01-01 08:58:27 +01:00