gemma.cpp

Commit Graph

Author	SHA1	Message	Date
Jan Wassenberg	ceb70203f0	Add min_verbosity to MaybePrint PiperOrigin-RevId: 886094998	2026-03-19 04:22:01 -07:00
Jan Wassenberg	529c201eb6	Add/use MaybePrint; also ShowConfig in non-interactive builds PiperOrigin-RevId: 882688835	2026-03-12 11:20:41 -07:00
Ray Smith	bea8b1cdbd	Replaced attention in ViT with flash - 8x speedup of image tokenizer on AMD PiperOrigin-RevId: 880877209	2026-03-09 08:46:04 -07:00
Krzysztof Rymski	029cfd0b33	Int8 + microscaling support for kv cache formats. Right now multiplication is done by converting to corresponding float format. Can yield up to 2x improvements for membw constrained shapes PiperOrigin-RevId: 880748493	2026-03-09 02:50:08 -07:00
Ray Smith	49cb438b1e	Rollback of erroneous rollback. PiperOrigin-RevId: 877376165	2026-03-02 06:50:26 -08:00
Jan Wassenberg	fbd44cee42	Fix Windows warnings PiperOrigin-RevId: 877338937	2026-03-02 04:53:25 -08:00
The gemma.cpp Authors	a3d994915f	No public description PiperOrigin-RevId: 877333188	2026-03-02 04:32:29 -08:00
Ray Smith	16c1b29b89	Rewrote flash attention to use BF16, transpose k and v, rewrote the task distribution, increase parallelism on decode, and use double the registers for the core of flash attention. PiperOrigin-RevId: 877308306	2026-03-02 03:11:01 -08:00
Jan Wassenberg	c6587efe70	Improve instrumentation for ViT parts PiperOrigin-RevId: 875302990	2026-02-25 13:10:44 -08:00
Jan Wassenberg	56fa6e4839	Internal change plus add U8 type, check MatPtrT type at compile time PiperOrigin-RevId: 867582875	2026-02-09 06:54:11 -08:00
Krzysztof Rymski	16a7ba2d6e	Internal changes PiperOrigin-RevId: 854171429	2026-01-09 06:35:36 -08:00
Jan Wassenberg	42e9cf557d	Internal change / remove unused PrintSpeed PiperOrigin-RevId: 853694463	2026-01-08 05:26:31 -08:00
Krzysztof Rymski	2ee1fac74c	Internal changes PiperOrigin-RevId: 853138600	2026-01-07 01:21:37 -08:00
Krzysztof Rymski	44dfd69b9b	Internal changes PiperOrigin-RevId: 844759322	2025-12-15 07:14:37 -08:00
Jan Wassenberg	0c64987a96	Abort if args are unrecognized, refactor argument passing This catches typos/incorrect usage. Refactor: group Loader/Threading/Inference into GemmaArgs. All *Args ctors now have an extra ConsumedArgs& argument. PiperOrigin-RevId: 844690553	2025-12-15 03:18:45 -08:00
Jan Wassenberg	73c3627b67	Add tensor stats and output tensor_info: add missing header io: fix mode weights.h: add layer_idx to LayerWeightsPtrs PiperOrigin-RevId: 843531051	2025-12-11 22:52:46 -08:00
Martin Stolle	bfc0dfcfca	Enable flags= parsing PiperOrigin-RevId: 843103750	2025-12-11 01:17:59 -08:00
Krzysztof Rymski	64178ace38	Internal changes PiperOrigin-RevId: 842727112	2025-12-10 07:55:17 -08:00
Jan Wassenberg	5a6895c609	Avoid warning when OS affinity limits us to the second socket Also simplify NumSMT, detect from .smt field directly PiperOrigin-RevId: 841749486	2025-12-08 07:10:43 -08:00
Jan Wassenberg	1564dd3111	Fix empty enabled_lps in topology detection Also expand the debug output. PiperOrigin-RevId: 838832605	2025-12-01 10:23:47 -08:00
Jan Wassenberg	3c9e6cf113	Expand debug output for topology PiperOrigin-RevId: 837738553	2025-11-28 00:19:33 -08:00
Jan Wassenberg	ccb49bc82f	Add ToFloatSlow, move RandomFloat to test_util PiperOrigin-RevId: 837412290	2025-11-27 00:14:51 -08:00
Jan Wassenberg	091b4567c9	Minor: ParallelismStrategy->Parallelism PiperOrigin-RevId: 828936578	2025-11-06 06:56:10 -08:00
Jan Wassenberg	a344a70c59	Change (old) attention behavior to disallow wraparound, enforced via assertion. Shared kU64PerLine constant PiperOrigin-RevId: 828072451	2025-11-04 11:52:40 -08:00
Jan Wassenberg	3cc0139ebb	Fix excessive KC/MC from prior change This could lead to stack overflow in B_storage. Also do not require specific type for query_norm_scale, update batch sizes for attention tensors, more verbose Mat shape/type checks. PiperOrigin-RevId: 824987689	2025-10-28 05:33:01 -07:00
Biruk Mammo	5a05857deb	[Gemma.cpp] Allows non-owned arguments for attention methods. * Adds and uses a new `AttentionActivationPtrs` that holds non-owning `MatPtrs`. Acts as a view into `AttentionActivations`. * Updates `QBatch` to hold non-owning `MatPtr`s to the kv caches. * Enables the `MatPtrT` default constructor for simpler initializations. * Pulls out and passes `LayerWeightsPtrs::query_norm_scale` directly. While `LayerWeightsPtrs` already held non-owning `MatPtr`s, this change avoids the need to find and construct several empty weight tensors just to construct one `query_norm_scale` tensor. PiperOrigin-RevId: 824584177	2025-10-27 10:43:25 -07:00
Jan Wassenberg	86200ce224	1.01x speedup: improved autotune Group M=4..7 into same config. Add configs for power of two sizes. Allow odd mc to enable a single range for odd M. io.cc: warning fix(cast). IsBlock -> !IsOneMC benchmark_helper: best for verbosity 3, all configs for 4 ops_test: remove unused includes PiperOrigin-RevId: 824475104	2025-10-27 05:35:31 -07:00
Jan Wassenberg	a48e614f64	1.02x speedup: improve load balance and simplify parallelFor Remove ParallelizeOne/TwoRange, use ParallelForAcross/WithinCluster instead. PiperOrigin-RevId: 823388890	2025-10-24 00:19:09 -07:00
Jan Wassenberg	3ed403e287	Major cleanup of profiler zones, add Caller annotation for all pool.Run Pass ThreadingContext instead of Pools/Profiler individually, for access to Zones Add GCPP_ZONE helper Add Caller argument to pool.Run to enable new stats Remove most direct dependencies on ThreadPool, prefer ParallelFor PiperOrigin-RevId: 822934530	2025-10-23 01:54:24 -07:00
Jan Wassenberg	acede9d682	Warning fix (unused var), Windows build fix (missing member variable) PiperOrigin-RevId: 822172982	2025-10-21 10:17:34 -07:00
Jan Wassenberg	f59eb2ed72	Remove multi-package support from topology Also no longer assume equal-sized clusters PiperOrigin-RevId: 820164125	2025-10-16 04:00:35 -07:00
Phil Culliton	503aaddd65	Add 8-bit integer quantization (I8Stream) to Gemma.cpp. PiperOrigin-RevId: 819787856	2025-10-15 09:25:20 -07:00
Ray Smith	e3e8511e79	Initialization of profiler zones. PiperOrigin-RevId: 819662587	2025-10-15 03:05:58 -07:00
Ray Smith	fb6fa793f4	Added a global (to gemma) zones list to enable most call sites to PROFILER_ZONE3 to avoid the sychronization required for the static const initialization of the zone handle. Improved flash_attention to enable profiling using the new zones. PiperOrigin-RevId: 819235421	2025-10-14 08:30:58 -07:00
Jan Wassenberg	035273c184	tune pool kSpin mode in threading_context Previously, this happened concurrently with the matmul autotune, which could lead to incorrect outcomes. threading: de-singleton Pinning (no longer stores affinity); pass PoolWorkerMapping; fix Pool dtor order Also enable SPR target (Zen4 is AMD-only), update Highway version for renamed Thread()->GlobalIdx(). PiperOrigin-RevId: 816223017	2025-10-07 08:36:26 -07:00
Jan Wassenberg	f3bc1c17da	1.03x speedup: fused FFN matmul-inl: support CView=StridedView or RowPtrs; rename to C_MC_NC matmul.cc: Allow 1 more rep for MC/NC to allow half-sized tiles, which helps. PiperOrigin-RevId: 807291701	2025-09-15 10:26:37 -07:00
Jan Wassenberg	ba6131311a	Fix gemma_batch_bench for flash attention q_T rows do not change. Also repeat prefill to reflect perf after autotuning. PiperOrigin-RevId: 805319377	2025-09-10 05:32:34 -07:00
Jan Wassenberg	9457258330	Refactor MatMul to accept views in the kernel functions Make arg order consistent. Move StridedView into mat.h. Add view support to RowPtrs. PiperOrigin-RevId: 805197381	2025-09-09 22:09:47 -07:00
Jan Wassenberg	24b1760f03	Refactor: move Worker to ThreadingContext, factor out MMDecompress PiperOrigin-RevId: 804909921	2025-09-09 07:56:12 -07:00
Jan Wassenberg	461a9c7d1b	Matmul refactoring towards fusion MMLoops: move dispatch code out, use overloads split build target into matmul_env (for MatMulEnv/MMOptions) weights: no longer call BindB Fix potential out of bounds in gemma_batch_bench PiperOrigin-RevId: 804895985	2025-09-09 07:13:38 -07:00
Jan Wassenberg	a5ab99e4ba	Memory use reduction: smaller/single MMStorage PiperOrigin-RevId: 804865029	2025-09-09 05:32:46 -07:00
Jan Wassenberg	06e5da1e22	Cleanup: split CacheInfo from Allocator, MatMul helper functions Lift DecompressA out of main autotuner to prevent interference Also use kMaxNR / kNR constants instead of extra args Fix: only require vector alignment, not cache alignment PiperOrigin-RevId: 804333769	2025-09-08 02:23:58 -07:00
Jan Wassenberg	56186193c1	Replace mt19937 with new generator to enable parallel sampling Split it into immutable AesCtrEngine and RngStream Also add RowSpan and Logits span PiperOrigin-RevId: 803336423	2025-09-04 23:49:10 -07:00
Jan Wassenberg	afd82376a5	Add AES-CTR RNG for parallel sampling (not yet used) PiperOrigin-RevId: 802991142	2025-09-04 05:58:42 -07:00
Jan Wassenberg	4be4799727	Remove kMaxPackages and per-package-related code matmul: remove kMaxClusters, dynamic allocation PiperOrigin-RevId: 802950348	2025-09-04 03:33:12 -07:00
Jan Wassenberg	7263ab8445	MatMul simplification, threading strategy improvements remove MatMul f32 special case (smaller code), types: Add u32/u64 for use by Activations move renamed ParallelismStrategy to threading_context so can pass ctx ensure worker index is unique across clusters matmul.h: const member functions for renamed policy classes (easier to call) PiperOrigin-RevId: 802848086	2025-09-03 21:45:07 -07:00
Jan Wassenberg	b7b3d353db	Simplify MatMul: remove F32 special case (build time) Also move kMaxM into separate kMaxBatchSize PiperOrigin-RevId: 802086590	2025-09-02 04:29:21 -07:00
Jan Wassenberg	1e3c853e80	Add ParallelFor wrapper function and one new mode Move ParallelismType from matmul.h to threading.h Replace SmallParallelFor with ParallelFor and the new mode PiperOrigin-RevId: 802038452	2025-09-02 01:40:09 -07:00
Jan Wassenberg	98ddc166db	Expand ThreadingContext comments PiperOrigin-RevId: 800479954	2025-08-28 08:32:10 -07:00
Jan Wassenberg	faa4102992	(Resubmit) Prepare profiler annotations for new API Pass hwy::Profiler& to low-level functions. Used ThreadingContext arg instead of NestedPools. Use new PROFILER_ZONE3. PiperOrigin-RevId: 794461159	2025-08-13 01:38:24 -07:00

1 2 3 4

182 Commits