gemma.cpp

Commit Graph

Author	SHA1	Message	Date
RangerUFO	7aac765e96	Add `Append` method to `AllQueries`	2025-06-16 20:39:27 +08:00
Jan Wassenberg	e5c81f64a1	Major refactor: clarify query_idx (global) vs qi. Refs #607 Fix missing pos increment for last prefill and check that in gemma_test. Thanks to @ufownl for pointing this out. Change argument lists to QBatch with accessors. Increase default seq_len to 8k. PiperOrigin-RevId: 771937385	2025-06-16 02:42:02 -07:00
Jan Wassenberg	2c72ff2aa5	Fix MatMul issue caused by autotuning bucketing, refs #608 , thanks @ufownl PiperOrigin-RevId: 771077158	2025-06-13 06:58:42 -07:00
Jan Wassenberg	01cdefeda7	1.64x batch=1 prefill speedup: nested parallelization for Attention (DotSoftmaxWeightedSum) Also fix tsan error in matmul (atomic_flag instead of static) PiperOrigin-RevId: 770241705	2025-06-11 11:28:46 -07:00
Jan Wassenberg	c027a45a2e	MatPtr-ify KV, shared div_seq_len, --seq_len flag PiperOrigin-RevId: 770194455	2025-06-11 09:49:38 -07:00
Jan Wassenberg	bd98b43cea	Rename RowPtr->StridedView, CRows->RowPtrs PiperOrigin-RevId: 770046362	2025-06-11 02:30:53 -07:00
Jan Wassenberg	b84149310b	Fix paligemma, update its test Must not pass image tokens to the EmbedMMToken used for text. Caught by next presubmit test. paligemma_test: move function bodies into class, regroup variables PiperOrigin-RevId: 770040014	2025-06-11 02:12:12 -07:00
Jan Wassenberg	ec02726cf7	6x large-batch, short-prompt prefill speedup Parallelize over queries instead of tokens introduce non_eos so we only iterate over not yet EOS queries; remove TokenStreamer. move RMSNormInplaceBatched out of Transformer to call the latter from prefill Consistent arg order. Fix gemma_test EOS handling which (caught by msan), remove from tokenizer.h Also add output to gemma_batch_bench, fix name PiperOrigin-RevId: 769676106	2025-06-10 09:56:20 -07:00
Daniel Keysers	d7b23d532a	Restructure internal initialization. PiperOrigin-RevId: 769507096	2025-06-10 01:25:31 -07:00
Rhett Stucki	824a95793c	Fix Image::WriteBinary() writing values to a file one at a time. PiperOrigin-RevId: 767955187	2025-06-06 00:48:09 -07:00
Jan Wassenberg	6ee628ba38	Further cleanup: separate MatMulEnv arg move row_ptrs into MatMulEnv Consistent arg order: layer, activations, kv_cache, env PiperOrigin-RevId: 767886386	2025-06-05 20:48:32 -07:00
Jan Wassenberg	e774ddbaaa	Github test: disable failing ubuntu-20.04 Also attempt to speed up bazel build. PiperOrigin-RevId: 767667520	2025-06-05 10:30:38 -07:00
Jan Wassenberg	0e2cab5187	Avoid warning about inability to map, unless explicitly requested PiperOrigin-RevId: 767633815	2025-06-05 09:10:08 -07:00
Jan Wassenberg	3a266c662c	Split gemma-inl into separate source files weights, mat: zero-initialize padding, required since the MatMul "avoid B decompress" optimization. PiperOrigin-RevId: 767562313	2025-06-05 05:36:44 -07:00
The gemma.cpp Authors	dd7d4a7717	Optimize Image::GetPatch() to copy rows instead of pixels at a time. PiperOrigin-RevId: 767436146	2025-06-04 22:31:08 -07:00
Copybara-Service	eff0213e88	Merge pull request #593 from ufownl:bugfix/dc2bf16 PiperOrigin-RevId: 767098675	2025-06-04 05:21:54 -07:00
RangerUFO	a82f8d5690	Fix compilation error on G++ 9.4	2025-06-04 17:39:37 +08:00
Jan Wassenberg	6897313080	3x speedup of EmbedImagePatches - GEMM, not GEMV. Required fixes to handling of non-vector aligned A. Also move row ptrs to MatMulEnv. PiperOrigin-RevId: 767029036	2025-06-04 01:18:52 -07:00
Daniel Keysers	9f74a1a098	Fix a problem in run_example.py PiperOrigin-RevId: 767017932	2025-06-04 00:42:57 -07:00
Jan Wassenberg	9efdcfd45c	1.07x batch decode speedup: more BF16 weights and activations BF16 att_sums and ffw_out Support BF16 B views without decompression Support arbitrary types in MulByConstAndAdd, AddFrom Also update profiler annotations in ops-inl.h PiperOrigin-RevId: 766995010	2025-06-03 23:30:18 -07:00
Jan Wassenberg	839a642992	Fix paligemma_test, refs #588 Detect PaliGemma models from layer names Remove unused allocator arg from CreateInvTimescale matmul: only warn once about dim divisibility Print config also in tests if --verbosity 2 PiperOrigin-RevId: 766605131	2025-06-03 04:45:22 -07:00
Copybara-Service	209009b57e	Merge pull request #588 from ufownl:bugfix/vit_attn PiperOrigin-RevId: 766528391	2025-06-03 00:43:30 -07:00
Jan Wassenberg	ad3002a21c	Merge branch 'dev' into bugfix/vit_attn	2025-06-03 09:29:52 +02:00
Jan Wassenberg	794a21a4e6	Major refactor to de-templatize gemma-inl and weights This replaces per-weight instantiations of all code with only per-MatMul/norm. Reduces binary size by 133KiB. WeightsOwner is no longer required for type erasing, hence it is replaced with ModelWeightsPtrs. Also remove unused EmbedToken, replaced with EmbedMMToken. PiperOrigin-RevId: 766497657	2025-06-02 23:01:35 -07:00
RangerUFO	93de2be938	Fix the broken VitAttention	2025-06-03 12:40:13 +08:00
Jan Wassenberg	cf4d7ceb82	1.16x decode speedup: remove last MatVec in Attention Precompute row pointers. Remove no longer used MHA support; QStride -> qkv_dim. Remove RowPtr from MatMul interface, use only MatPtrT. Require opt-in define for NUQ to speed up builds. Also fix io.cc on Windows. PiperOrigin-RevId: 766228108	2025-06-02 09:40:29 -07:00
Jan Wassenberg	c4a75abe43	Cleanup gemma_batch_bench PiperOrigin-RevId: 766177406	2025-06-02 07:04:36 -07:00
Jan Wassenberg	a3f7bf0991	Fix thread name when skipping packages/clusters PiperOrigin-RevId: 766054198	2025-06-01 23:50:11 -07:00
Jan Wassenberg	0023ff8770	Add support for arbitrary output row pointers Useful for writing directly to KV cache. PiperOrigin-RevId: 765615147	2025-05-31 10:55:54 -07:00
The gemma.cpp Authors	9c3e089b09	Internal change. PiperOrigin-RevId: 765218260	2025-05-30 09:18:44 -07:00
The gemma.cpp Authors	1e8642f8f4	Internal change. PiperOrigin-RevId: 765037449	2025-05-29 22:51:16 -07:00
Jan Wassenberg	3890eb5412	Remove backprop/ Also remove MatPtrT::Packed(); use PackedScale1 instead where const, or Row(0). PiperOrigin-RevId: 764243198	2025-05-28 07:01:17 -07:00
Jan Wassenberg	627cc04db9	Decouple MatMul from gemma-inl: precompile for all input types Call MatMulStatic instead of MatMul. Also fix build error due to Highway's Lanes not being constexpr. PiperOrigin-RevId: 763777269	2025-05-27 07:08:58 -07:00
Jan Wassenberg	421a2ab8ac	Add comments explaining non-padded tensors, kNoPad -> kPacked PiperOrigin-RevId: 763352173	2025-05-26 03:03:38 -07:00
Copybara-Service	eb8a463038	Merge pull request #574 from ufownl:bugfix/vit_weights PiperOrigin-RevId: 761948356	2025-05-22 07:04:53 -07:00
RangerUFO	2771f463f9	Fix the ViT weights loading	2025-05-22 12:13:29 +08:00
Copybara-Service	1ce89788ef	Merge pull request #573 from ufownl:bugfix/vit PiperOrigin-RevId: 761425663	2025-05-21 01:58:00 -07:00
RangerUFO	6debdbe341	Minor fixes for ViT	2025-05-20 22:27:10 +08:00
Jan Wassenberg	cb188d4a0e	Fix RowT issue and improve Griffin (currently still broken) Use type-safe MatPtrT via dynamic_cast, avoid/remove unsafe RowT activations: Griffin tensors are now padded Griffin: add batching support, fix conv1d_cache allocation weights: bundle to TensorToRead, add kNoPad flag, fix SplitW1 const-correct fix for ForEachTensor blob_store: move BlobIO2 to .cc and rename BlobIO PiperOrigin-RevId: 760610094	2025-05-19 07:02:10 -07:00
Jan Wassenberg	d6cfabc2c1	Shorten gemma_test so we can run it for more models. PiperOrigin-RevId: 759685282	2025-05-16 11:14:41 -07:00
Jan Wassenberg	e890d46f30	1.31x batch prefill, 1.24x batch decode speedup: NUMA binding Only the weights; binding MatMul output worsens batch=1 prefill. Update gemma_batch_bench to use --decode_qbatch. Fix/remove prefill_activations in gemma-inl.h. Refactor: use BasePageBytes directly when binding Move BindB/C to .cc by de-templatizing Remove MatOwners::AllocateFor because it is weights-specific (binding or not) Disband MatOwners, replace with vector PiperOrigin-RevId: 759610477	2025-05-16 07:42:13 -07:00
Jan Wassenberg	c443adee33	3.8x speedup of weights loading via preadv on Linux Also move BlobReader reading functionality to weights.cc PiperOrigin-RevId: 759240310	2025-05-15 11:55:15 -07:00
Jan Wassenberg	38a08d8095	Replace last ConstMat with MatPtr This is to reduce the number of MatMul overloads in preparation for de-templatizing. PiperOrigin-RevId: 758288589	2025-05-13 10:55:22 -07:00
Copybara-Service	0a6a7e4cd6	Merge pull request #566 from ufownl:bugfix/deduced_model_wrapping PiperOrigin-RevId: 758276145	2025-05-13 10:28:16 -07:00
RangerUFO	30ad625f42	Fix the wrapping field of the deduced model config	2025-05-13 23:02:03 +08:00
Jan Wassenberg	8a312e9b89	Split W1/W2 as a load-time preprocess. Remove kOnlyAllocate - no longer used. Rename ReadOrAllocate -> ReadFromBlobs. Rename Reshape -> Fixup to reflect the new scope. Remove no longer used ShrinkRows. This simplifies gemma-inl and is a prerequisite for removing ConstMat (whose .ofs was previously used for merged tensors) PiperOrigin-RevId: 758214083	2025-05-13 07:39:59 -07:00
Jan Wassenberg	2038dfd9cc	Minor: rename compression/shared -> types.h PiperOrigin-RevId: 758199851	2025-05-13 06:53:21 -07:00
Jan Wassenberg	d538a6d6c6	Cleanup: remove unused kCyclic, remove 2 suffix Also remove now unused allocator arg and fix warnings (cast, struct/class mismatch) PiperOrigin-RevId: 758098495	2025-05-13 01:06:41 -07:00
Biruk Mammo	ba21e3beb4	Adds a `GemmaAttention` constructor that takes an explicit `ThreadingContext`. PiperOrigin-RevId: 757839682	2025-05-12 11:17:05 -07:00
Jan Wassenberg	45ad847a41	Replace RowVectorBatch with MatStorageT KVCache: add ctor required for MatStorageT, remove Create; bf_pre_ffw_rms_out -> pre_ffw_rms_out optimize_test: larger vocab_size requires more steps shared.h: Remove unused u128 type correctly set Activation matrix rows, avoid passing as arg ops: pass Mat instead of pointers/sizes; vectorize LayerNorm; support any weight type mat: add OverrideRows, used by SetBatchSize PiperOrigin-RevId: 757790736	2025-05-12 09:16:12 -07:00

1 2 3 4 5 ...

674 Commits All Branches Search

674 Commits

All Branches