llama.cpp

Commit Graph

Author	SHA1	Message	Date
HanishKVC	7302b3ab36	SimpCfg: Use stderr wrt internal Log messaging helpers	2024-05-06 11:27:56 +05:30
HanishKVC	a09571318a	ChatON: meta-dump returns flag inturn returned by meta-ok test-chat-template-chaton now tries to check if meta-ok is ok wrt the template-id being looked into. Log template-id info also, where it was previously missed out.	2024-05-06 11:27:56 +05:30
HanishKVC	44c05305d0	SimpCfg: Add support for get_double	2024-05-06 11:27:56 +05:30
HanishKVC	8ad2c17e5d	SimpCfg: get_int64 logic	2024-05-06 11:27:56 +05:30
HanishKVC	000245b8e8	SimpCfg:Warn possible nonstring strings, some invalid floats Warn if something not starting with double quote is being treated as a string. Show some examples of invalid floating point values wrt this logics floating point determination code	2024-05-06 11:27:56 +05:30
HanishKVC	a6648b02f2	SimpCfg:Show floating point values in normal and exponential form	2024-05-06 11:27:56 +05:30
HanishKVC	4181164217	SimpCfg:Implement set_int64 and set_double Also update the sample simpcfg file, to test for int and float values.	2024-05-06 11:27:56 +05:30
HanishKVC	fb9a7dc7fe	SimpCfg:Initial skeleton towards supporting int and floating point	2024-05-06 11:27:56 +05:30
HanishKVC	d0b3ebf32e	SimpCfg: Use & wrt destination of [] operation so that one can update the value-item's content, without needing to explicitly update/store the value-item back into map after the content has been updated. This should make these setting operations/helpers more efficient.	2024-05-06 11:27:56 +05:30
HanishKVC	0a534e6897	SimpCfg: Rename test program related #define	2024-05-06 11:27:56 +05:30
HanishKVC	ca5a04d607	SimpCfg: Remove double quotes around group, key or string value	2024-05-06 11:27:56 +05:30
HanishKVC	82348e2840	SimpCfg: Put GroupMap back into Map, Iterate during get if DBUG TODO: Have to look into C++ a bit later including these default container types. Not sure my current flow is efficient.	2024-05-06 11:27:56 +05:30
HanishKVC	9940bd8ed7	SimpCfg: Allow default values wrt set string and set bool	2024-05-06 11:27:56 +05:30
HanishKVC	951fbc3396	SimpCfg: Change logging to LDBUG and LERRR helpers	2024-05-06 11:27:56 +05:30
HanishKVC	d514c81829	SimpCfg: Add the const which I had forgotten wrt args	2024-05-06 11:27:56 +05:30
HanishKVC	6de8a14f32	SimpCfg: Rename member functions to avoid sc_ prefix now that logic has been converted into a class, no need for this prefix	2024-05-06 11:27:56 +05:30
HanishKVC	1ecca5a7ec	SimpCfg: Convert to a class	2024-05-06 11:27:56 +05:30
HanishKVC	28ae0c5b02	SimpCfg:Make str_trim flexible, use to trim , wrt value Now one can pass the list/string of chars to trim at either end.	2024-05-06 11:27:56 +05:30
HanishKVC	2cbb00c340	SimpCfg: Add support for boolean fields wrt key-value	2024-05-06 11:27:56 +05:30
HanishKVC	aea6850131	SimpCfg: Keep compiler happy, also add newline wrt alt logging def	2024-05-06 11:27:56 +05:30
HanishKVC	f4687fa5d4	SimpCfg:Parse config file and load string key-value fields	2024-05-06 11:27:56 +05:30
HanishKVC	ce75d434dc	SimpCfg: Initial skeleton : get and set string and bool values	2024-05-06 11:27:56 +05:30
HanishKVC	af9a0a211b	ChatON:ChatTmplApply: Avoid the stringstream	2024-05-06 11:27:56 +05:30
HanishKVC	889a45ff28	ChatON:ChatTmplApply:Update the function notes	2024-05-06 11:27:56 +05:30
HanishKVC	ff5f68826b	ChatON:ChatTmplApplySingle: Avoid streamstring, update func notes	2024-05-06 11:27:56 +05:30
HanishKVC	32e672c5dd	ChatON: Dont log final tagged message string to screen	2024-05-06 11:27:56 +05:30
HanishKVC	cad50c527e	ChatON: Update the note to match current logic	2024-05-06 11:27:56 +05:30
HanishKVC	a4b3285034	ChatON:Show Log on screen when template is applied	2024-05-06 11:27:56 +05:30
HanishKVC	d61b071b8d	Chaton:Common:Add missing newline wrt cmdline arg usage	2024-05-06 11:27:56 +05:30
HanishKVC	fee887fe31	ChatON:Common:Update the cmdline argument name used Had forgotten to update it before	2024-05-06 11:27:56 +05:30
HanishKVC	58e1ff16bc	ChatON: switch to ordered_json from json library to be in sync with the json namespace in server.	2024-05-06 11:27:56 +05:30
HanishKVC	a630564c48	ChatON:ChatTemplateApplyCAPI remaining base logic As c doesnt have the concept of pass by reference, and inturn the existing c api uses pointers wrt llama chat message structure, so switching to same wrt chat_tmpl_apply logics. Also fix a oversight in previous commit and add the remaining logic.	2024-05-06 11:27:56 +05:30
HanishKVC	308d3bf3ff	ChatON:WIP:Add c api wrapper for chat_template_apply Initial skeletons Update existing logics to help with same. Also the inbetween helper was having a bad signature wrt returning status and data, thats also fixed.	2024-05-06 11:27:56 +05:30
HanishKVC	e62699f923	ChatON: Add alertAssistantAtEnd flag & logic wrt MultiMsgs Apply While sending the current chat session along with new user query to the model, many models expect that a tag be added at the end to indicate that user is expecting the model to respond, this flags allows for the same.	2024-05-06 11:27:56 +05:30
HanishKVC	ea3a0f19cc	ChatON: Rather check for tmpl existance in single_ex	2024-05-06 11:27:56 +05:30
HanishKVC	01c8db70f7	ChatON+Main: Add C_API wrapper for single Add a c api wrapper for a single message tagging scenario. Inturn to match convention followed by existing chat_apply_template code, make it return the size expected of the tagged message string buffer. Update internal single logic to help with same. Explicitly check if tmpl specified is available in the loaded json or not and then return a error if not found.	2024-05-06 11:27:56 +05:30
HanishKVC	13857f29d6	ChatON+Main: Updates wrt detailed meta json Fix a oversight wrt key name. Add a alert in case if passed meta json file contains begin(BoS) wrt assistant role, similar to check for end (EoS) wrt user role. Bcas normally both (ie EoS wrt User and BoS wrt Assistant) shouldnt be needed. Update main wrt begin & prefix and suffix & end addition.	2024-05-06 11:27:56 +05:30
HanishKVC	0cd7c62706	ChatON: Keep compiler happy Move helpers to the begining, so can avoid adding prototype declerations/function signatures to the begining Get the char * wrt string data in the c++ string.	2024-05-06 11:27:56 +05:30
HanishKVC	6a0214c067	ChatON:MetaOK->MetaDump: Alert if user->end is needed or not Because user messages dont normally need a EoS token.	2024-05-06 11:27:56 +05:30
HanishKVC	344857b6cb	ChatOn:ChatOnTemplateApply: suffix,end flag based control Also fix a oversight wrt begin, when flag based begin adding control was introduced. NOTE: Currently system role suffix/end conditional adding always triggered, if 1st system prompt seen or additional system prompt is seen.	2024-05-06 11:27:56 +05:30
HanishKVC	f8ae21cec7	ChatON:ChatTemplateApplySingle: update begin+prefix, suffix+end	2024-05-06 11:27:56 +05:30
HanishKVC	5d76f08d37	ChatON: Need to explicitly specify string to use c_str	2024-05-06 11:27:56 +05:30
HanishKVC	7ba0144e42	ChatOn:chaton_tmpl_role_kv: try except to ignore missing ifany Cas of above reason, switch to directly accessing the keys in dump helper, which is inturn used by meta_ok check	2024-05-06 11:27:56 +05:30
HanishKVC	adab5775bf	ChatON: more detailed/spreadout json fields	2024-05-06 11:27:56 +05:30
HanishKVC	3f09eb5dea	ChatOn: ChatTemplateApply[Ex] return tagged msgs parts detail Now there is a simple and extended version of returning tagged messages. The extended version returns the tagged string, as well as the details of the parts that make up that tagged message interms of the type of parts and the lengths of the parts.	2024-05-06 11:27:56 +05:30
HanishKVC	825a78abaa	ChatOn: ChatTemplateApplySingle[Ex] return parts detail Now there is a simple and extended version of returning tagged message wrt a single role and its content. The extended version returns the tagged string, as well as the details of the parts that make up that tagged message interms of the type of parts and the lengths of the parts.	2024-05-06 11:27:56 +05:30
HanishKVC	92e780fb1a	ChatON:ChatParts: Allow flexibility for more refined tokenization	2024-05-06 11:27:56 +05:30
HanishKVC	d1899728aa	ChatON: Test ChatParts in chat-template-apply	2024-05-06 11:27:56 +05:30
HanishKVC	9de1d6017f	ChatON:ChatParts class initial go Helps keep user prompt and chat-hs-template tag parts seperate, but in sequence	2024-05-06 11:27:56 +05:30
HanishKVC	3064a36e74	ChatON+:Update tmpl_role_kv to retrieve wrt multiple keys Use the same for user role's begin and prefix entries.	2024-05-06 11:27:56 +05:30
HanishKVC	f1f39c5256	ChatON:Add Monarch model template, which uses Begin + Prefix Inturn Begin/BoS is added only for non 1st user messages in a system+user prompts chain.	2024-05-06 11:27:56 +05:30
HanishKVC	724ff38345	ChatOn: Wrap getting begin in try-catch, so that even if a role doesnt contain begin, the logic will work fine.	2024-05-06 11:27:56 +05:30
HanishKVC	d70fca7a45	ChatOn: Add begin to the mix along with prefix Dump shows user->begin. chat-template-apply[-single] updated to work with begin and prefix TODO: need to wrap begin in a try-catch, so that irrespective of role, begin+prefix will work, irrespoective of whether that role has a begin entry or not.	2024-05-06 11:27:56 +05:30
HanishKVC	bdd279c0c9	ChatOn:User Begin+Prefix note update, keep things simple consistent	2024-05-06 11:27:56 +05:30
HanishKVC	84367b9fd1	ChatON: Add template for DeepSeek Was looking at the tokenized vector, and noticed that the EOS mentioned by existing chat_apply_template of llama.cpp, is different from what I noticed in tokenizer_config.json of deepseek llm, so I have added two entries * "deepseek-alt" which matches llama.cpp's chat_apply_template and * "deepseek" which matches that in tokenizer_config.json. This impacts the assistant suffix and reverse prompt entries. CasOfThis: Need to look into other entries which I added previously at a later time. However as the default logic should be picking the EOS from model file, so I assume reverse-prompt being outofsync, may not matter beyond a limit, potentially.	2024-05-06 11:27:56 +05:30
HanishKVC	57bd772bfd	ChatON: Cleanup logging Avoid showing on screen the debug messages. meta-dump can either show on screen or not, based on how LOGXLN is defined.	2024-05-06 11:27:56 +05:30
HanishKVC	217544e5ff	ChatON: Keep compiler happy Order the functions so that no need for seperate prototypes Also use kv_bool wrt boolean entries. Convert string to c char *	2024-05-06 11:27:56 +05:30
HanishKVC	3f9dfc240c	ChatON: Check for the boolean entries in meta-json	2024-05-06 11:27:56 +05:30
HanishKVC	42f6b45547	ChatON: Use the constants defined for the keys	2024-05-06 11:27:56 +05:30
HanishKVC	efb758ba7d	ChatON: Rename helpers to kv suffix, updated wrt metaok rename because they return value of specified key. [main] update metaok to take template-id, so that one can cross check that all needed entries are there wrt that template-id in the chaton-meta-json file	2024-05-06 11:27:56 +05:30
HanishKVC	e8c24c0767	ChatOn:MetaOk: Allows template-id based cross check For a given template-id, cross check, all needed entries are there in the json.	2024-05-06 11:27:56 +05:30
HanishKVC	b1055641e9	ChatON: Update the notes a bit	2024-05-06 11:27:56 +05:30
HanishKVC	11b47fbcfc	ChatON:MetaJson: Add key constants, check metaJson loaded ifNeeded	2024-05-06 11:27:56 +05:30
HanishKVC	221ccd6462	ChatOn: Add SystemUser-1st-User-Has-Prefix flag support Llama2 seems to need it, so chaton-meta-json sample file updated to use same.	2024-05-06 11:27:56 +05:30
HanishKVC	f03dd2439f	ChatOn:No global-begin/end in ChatApplyTmplSingle, ChatApplyTmpl Avoid adding global begin/end markers wrt ChatApplyTmplSingle. Add ChatApplyTmpl which goes through a vector of messages.	2024-05-06 11:27:56 +05:30
HanishKVC	c4cf0e9075	ChatON:Cleanup: BeginEnd, Debug log Update the note Rename global-prefix\|suffix to global-begin\|end. Rename chat-apply-template to chat-apply-template-single, cas it handles only a single message. Add some debug log messages to the helper functions	2024-05-06 11:27:56 +05:30
HanishKVC	050d329e7e	ChatOn+Main: Initial go at chaton in main interactive flow	2024-05-06 11:27:55 +05:30
HanishKVC	dc56be951d	ChatOn:Main: Load and dump any specified chaton meta file	2024-05-06 11:27:55 +05:30
HanishKVC	35f25196a0	ChatOn:Common: Add the needed cmdline arg params and its parsing	2024-05-06 11:27:55 +05:30
HanishKVC	2146a253e8	ChatOn: Capture the idea	2024-05-06 11:27:55 +05:30
viric	fcd84a0f5a	Fix Linux /sys cpu path to guess number of cores (#7064 )	2024-05-04 15:26:53 +02:00
Andrew Downing	b0d943de17	Update LOG_IMPL and LOG_TEE_IMPL (#7029 ) ROCm clang defines _MSC_VER which results in the wrong implementation of LOG_IMPL and LOG_TEE_IMPL being compiled. This fixes https://github.com/ggerganov/llama.cpp/issues/6972	2024-05-01 23:31:30 +02:00
Johannes Gäßler	a8f9b07631	perplexity: more statistics, added documentation (#6936 ) * perplexity: more statistics, added documentation * add LLaMA 3 8b scoreboard	2024-04-30 23:36:27 +02:00
Georgi Gerganov	9c67c2773d	ggml : add Flash Attention (#5021 ) * ggml : add ggml_flash_attn_ext API * ggml : fix GQA support in ggml_flash_attn_ext * ggml : online attention (CPU) * metal : initial implementation * metal : f16 precision * metal : reduce branches * metal : specialize for head size * wip : 8 rows per simd group * wip : 4 rows per simd group * wip : template for rows per warp * metal : parallelize across KV size * metal : parallel reduce across heads * metal : efficient flash_attn_f16 implementation * metal : avoid redundant loads of the attention * metal : scale and mask in matrix form * metal : fix comment * llama : avoid ggml_cast, use F32 query * metal : add parallel reduce version (disabled) * metal : move output into local memory + optimize - the result from each simdgroup now stays in the registers - significantly reduced SRAM usage - more efficient skipping of -INF blocks - avoid simdgroup barrier in hot loop - add comments * metal : add tests, fix scaling, support C > 32 * metal : improve precision * ggml : fix f16 mad * metal : minor * metal : support Q > 8 * tests : add ATTN tests * metal : disable buffer allocation logs * tests : more * metal : faster inner loop for C == 32 * metal : fix array initialization * tests : ifdef * ggml : switch to padded F16 mask for ggml_soft_max, ggml_flash_attn_ext * ggml : fix ggml_soft_max mask requirement * cuda : fix soft_max to use correct mask size * cuda : add flash_attn kernel (wip) * metal : optimize softmax for C > 32 * metal : optimize softmax * tests : minor fix * cuda : avoid zeroing fragments * tests : update dims * cuda : fix __hisinf() result check * cuda : avoid warp_reduce for smax * cuda : use int instead of int64_t Noticeably improves performance (thanks to Johannes) * cuda : make loops use the same loop values Thanks Johannes again for the tip * cuda : unroll some of the loops * cuda : avoid __hisinf branches * cuda : use half2 in softmax * cuda : switch to 1 warp for bs > 16 * cuda : speed-up reduce part of the kernel * cuda : unroll QK^T loop cuda : fix -INF block check * cuda : simplify softmax * cuda : fix matrix names * cuda : minor * llama : adapt to F16 KQ_pos * llama : adapt new models to F16 KQ_mask * ggml : fix F16 store (ARM NEON) * llama : fix type of KQ_mask and KQ_pos * ggml : fix CPU soft_max * tests : add hs=256 * cuda : fix build * metal : improve perf via smaller int registers * cuda : adapt soft_max to F16 mask and pos * CUDA: faster FlashAttention, kernel for bs == 1 * 16 cols for Phi-2 * no vec for hs, no hs==256 ncols==32 for Volta * adjust kernel selection logic * 4 warps, 256 stride for all D * no ncols == 64 * Multiple parallel blocks for batch size 1 * fix compile warnings * fix excessive KQ_b loads * fix cmake build * fix KV cache padding, NaN from INFINITY (#6438) * llama : flash_attn cparam + fix defrag * server: support flash_attn param * server: bench: enable flash_attn param * CUDA: refactor host code, dyn. par. blocks * fix flash_attn_vec_f16 race condition * flush softmax exp below threshold to 0 * store temp KQ in registers * Calculate KQ as FP32 if KQV has GGML_PREC_F32 * Add __hgt2_mask implementation for CUDA 11 * fix KQ FP32 precision fpr parallel_blocks > 1 * llama-bench : add -fa,--flash-attn arg * metal : add BS=1 kernel for flash attention (#6508) * metal : add BS=1 kernel for flash attention (wip) * metal : support more than 1 warps * metal : opts * metal : opt * metal : switch to parallel reduce * metal : reduce registers * metal : simplify * metal : initial FA vec kernel * metal : use F32 attention accumulators * batched-bench : add fattn arg * llama : simplify llama_build_kv_store ggml-ci * llama : adapt build_olmo to changes * ggml : fix arm fp16 store on windows * metal : clean-up * metal : clean-up kernel code * metal : minor * tests : remove benchmarks ggml-ci * ggml : fix avx512 const correctness ggml-ci * ggml : fix soft_max with bias on CPU ggml-ci * common : print --flash-attn in help * ggml : fix num dimensions in ggml_flash_attn_ext * llama : force disable flash attention for incompatible models * ggml : ggml_soft_max support F16/F32 mask/pos ggml-ci * cuda : uint -> uint32_t * cuda : "constexpr dim3" -> "const dim3" ggml-ci * cuda : try to fix __hgt2_mask ggml-ci * ggml : add TODO's for F16/F32 mask/pos support in other backends * llama : replace bool need_kq_pos with use_alibi * llama : prep ALiBi support for BERT models ggml-ci * llama : fix n_batch requirements ggml-ci * cont * server : add help for --flash-attn arg * llama : disable FA for AMD * tests : remove TMP_ATTN_BENCH ggml-ci * llama : support save/load state with FA enabled ggml-ci * ci : add CUDA save-load-state tests ggml-ci * llama : llama_kv_cache_clear zeroes data + fix save-load seq ggml-ci * llama : fix copy-paste errors, add TODO * llama : disallow incompatible states * llama : update llama_state_get_size after v_trans field * metal : remove tmp log * llama : add static reminder for llama_state_get_size * metal : fix max nsg ggml-ci * ci : fix arg order ggml-ci --------- Co-authored-by: Johannes Gäßler <johannesg@5d6.de> Co-authored-by: Pierrick HYMBERT <pierrick.hymbert@gmail.com>	2024-04-30 12:16:08 +03:00
Olivier Chafik	8843a98c2b	Improve usability of --model-url & related flags (#6930 ) * args: default --model to models/ + filename from --model-url or --hf-file (or else legacy models/7B/ggml-model-f16.gguf) * args: main & server now call gpt_params_handle_model_default * args: define DEFAULT_MODEL_PATH + update cli docs * curl: check url of previous download (.json metadata w/ url, etag & lastModified) * args: fix update to quantize-stats.cpp * curl: support legacy .etag / .lastModified companion files * curl: rm legacy .etag file support * curl: reuse regex across headers callback calls * curl: unique_ptr to manage lifecycle of curl & outfile * curl: nit: no need for multiline regex flag * curl: update failed test (model file collision) + gitignore *.gguf.json	2024-04-30 00:52:50 +01:00
cpumaxx	ffe666572f	llava-cli : multiple images (#6969 ) Co-authored-by: root <root@nenya.lothlorien.ca>	2024-04-29 17:34:24 +03:00
Georgi Gerganov	f4ab2a4147	llama : fix BPE pre-tokenization (#6920 ) * merged the changes from deepseeker models to main branch * Moved regex patterns to unicode.cpp and updated unicode.h * Moved header files * Resolved issues * added and refactored unicode_regex_split and related functions * Updated/merged the deepseek coder pr * Refactored code * Adding unicode regex mappings * Adding unicode regex function * Added needed functionality, testing remains * Fixed issues * Fixed issue with gpt2 regex custom preprocessor * unicode : fix? unicode_wstring_to_utf8 * lint : fix whitespaces * tests : add tokenizer tests for numbers * unicode : remove redundant headers * tests : remove and rename tokenizer test scripts * tests : add sample usage * gguf-py : reader prints warnings on duplicate keys * llama : towards llama3 tokenization support (wip) * unicode : shot in the dark to fix tests on Windows * unicode : first try custom implementations * convert : add "tokenizer.ggml.pre" GGUF KV (wip) * llama : use new pre-tokenizer type * convert : fix pre-tokenizer type writing * lint : fix * make : add test-tokenizer-0-llama-v3 * wip * models : add llama v3 vocab file * llama : adapt punctuation regex + add llama 3 regex * minor * unicode : set bomb * unicode : set bomb * unicode : always use std::wregex * unicode : support \p{N}, \p{L} and \p{P} natively * unicode : try fix windows * unicode : category support via std::regex * unicode : clean-up * unicode : simplify * convert : add convert-hf-to-gguf-update.py ggml-ci * lint : update * convert : add falcon ggml-ci * unicode : normalize signatures * lint : fix * lint : fix * convert : remove unused functions * convert : add comments * convert : exercise contractions ggml-ci * lint : fix * cmake : refactor test targets * tests : refactor vocab tests ggml-ci * tests : add more vocabs and tests ggml-ci * unicode : cleanup * scripts : ignore new update script in check-requirements.sh * models : add phi-3, mpt, gpt-2, starcoder * tests : disable obsolete ggml-ci * tests : use faster bpe test ggml-ci * llama : more prominent warning for old BPE models * tests : disable test-tokenizer-1-bpe due to slowness ggml-ci --------- Co-authored-by: Jaggzh <jaggz.h@gmail.com> Co-authored-by: Kazim Abrar Mahi <kazimabrarmahi135@gmail.com>	2024-04-29 16:58:41 +03:00
David Renshaw	3f167476b1	sampling : use std::random_device{}() for default random seed (#6962 )	2024-04-29 16:35:45 +03:00
mgroeber9110	4dba7e8114	Replace "alternative" boolean operator in conditional compilation directive (#6949 )	2024-04-27 21:02:06 +02:00
Pierrick Hymbert	0c4d489e29	quantize: add imatrix and dataset metadata in GGUF (#6658 ) * imatrix: save the dataset file used in the output file * llama: support kv overrides type string string * common: factorize KV Overrides parsing between common and server * quantize: add imatrix n entries and dataset KV metadata quantize: factorize KV Overrides parsing between common #6656 * llama: remove kv override str_value initialization as it does not compile on some toolchain * quantize: add imatrix m_last_call as `quantize.imatrix.chunks_count` * quantize: add imatrix filename in KV * llama: add llama_model_kv_override_free * common: add llama_model_kv_override_free common: free kv override if used after model loading * llama: finally move the string KV override value to the stack * llama : minor * no need to add a NUL to the std::vector, std::string can be initialized from a pair of iterators. Co-authored-by: slaren <slarengh@gmail.com> * kv override: ensure string termination --------- Co-authored-by: Georgi Gerganov <ggerganov@gmail.com> Co-authored-by: slaren <slarengh@gmail.com>	2024-04-26 20:06:33 +02:00
slaren	017e6999b5	add basic tensor data validation function (#6884 ) * add basic tensor data validation function * add --check-tensors command line argument tensor validation is disabled by default and can be enabled by adding `--check-tensors` to the command line arguments. quantize always validates tensors.	2024-04-26 18:39:58 +02:00
Douglas Hanley	b4e4b8a935	llama : add llama_get_pooling_type function (#6862 ) * add llama_get_pooling_type function * fix argument name, move with ctx funcs	2024-04-24 16:10:07 +03:00
Kyle Mistele	37246b1031	common : revert showing control tokens by default for server (#6860 ) * fix: revert showing control tokens by default * feat: revert changes to default behavior of llama_token_to_piece; provide overridden declaration to receive "bool special" param to toggle showing control tokens * feat: use the overridden declaration of llama_token_to_piece from common/common.cpp to specify "false" so that control tokens are not shown in chat completion responses" * common : simplify --------- Co-authored-by: Georgi Gerganov <ggerganov@gmail.com>	2024-04-24 13:15:29 +03:00
Johannes Gäßler	28103f4832	Server: fix seed for multiple slots (#6835 ) * Server: add tests for consistent results * sampling: separate rng per sampling context	2024-04-24 11:08:36 +02:00
Georgi Gerganov	40f74e4d73	llama : add option to render special/control tokens (#6807 ) * make : fix common dep on llama.h * llama : add option to render special tokens * readme : add API change notice ggml-ci * swift : fix build	2024-04-21 18:36:45 +03:00
Georgi Gerganov	aed82f6837	common : try to fix Android CI (#6780 ) * common : disable get_math_cpu_count() until Android CI gets fixed * common : another try	2024-04-20 13:27:12 +03:00
Justine Tunney	8cc91dc63c	ggml : add llamafile sgemm (#6414 ) This change upstreams llamafile's cpu matrix multiplication kernels which improve image and prompt evaluation speed. For starters, Q4_0 and Q8_0 weights should go ~40% faster on CPU. The biggest benefits are with data types like f16 / f32, which process prompts 2x faster thus making them faster than quantized data types for prompt evals. This change also introduces bona fide AVX512 support since tinyBLAS is able to exploit the larger register file. For example, on my CPU llama.cpp llava-cli processes an image prompt at 305 tokens/second, using the Q4_K and Q4_0 types, which has always been faster than if we used f16 LLaVA weights, which at HEAD go 188 tokens/second. With this change, f16 LLaVA performance leap frogs to 464 tokens/second. On Intel Core i9-14900K this change improves F16 prompt perf by 5x. For example, using llama.cpp at HEAD with Mistral 7b f16 to process a 215 token prompt will go 13 tok/sec. This change has fixes making it go 52 tok/sec. It's mostly thanks to my vectorized outer product kernels but also because I added support for correctly counting the number of cores on Alderlake, so the default thread count discounts Intel's new efficiency cores. Only Linux right now can count cores. This work was sponsored by Mozilla who's given permission to change the license of this code from Apache 2.0 to MIT. To read more about what's improved, and how it works, see: https://justine.lol/matmul/	2024-04-16 21:55:30 +03:00
Olivier Chafik	7593639ce3	`main`: add --json-schema / -j flag (#6659 ) * main: add --json-schema / -j * json: move json-schema-to-grammar to common lib * json: fix zig build	2024-04-15 18:35:21 +01:00
Olivier Chafik	ab9a3240a9	JSON schema conversion: ⚡️ faster repetitions, min/maxLength for strings, cap number length (#6555 ) * json: rename python schema converter to make import easier * server: skip null json_schema / grammar fields * json: deps management for primitive rules (+ allow null values) * json: optimize repetitions for minItems/maxItems and regexps: `a{,3}` goes from `"a"? "a"? "a"?` (explosive combos) to `(a (a (a)?)?)?` * grammars: add troubleshooting section to readme * json: cap length of numbers to 15 digits before/after decimal point (avoids infinite gen, e.g. "one third" -> `0.333333333333...`) * json: unify all repetition code (w/ or w/o sep) * json: support string minLength/maxLength * server+json: update server/README w/ result_format * nits * json: fix type error w/ python 3.8 * json: fix server/README (json_schema in /completion vs. result_format in /v1/chat/completions) * json: simplify DOT `{"type": "string", "pattern": "^.$"}` * json: remove recursion in opt_repetitions (avoids Python stack overflow) * json: rm dead code * json: rm useless assert & ggml.h import	2024-04-12 19:43:38 +01:00
Pierrick Hymbert	b804b1ef77	eval-callback: Example how to use eval callback for debugging (#6576 ) * gguf-debug: Example how to use ggml callback for debugging * gguf-debug: no mutex, verify type, fix stride. * llama: cv eval: move cb eval field in common gpt_params * ggml_debug: use common gpt_params to pass cb eval. Fix get tensor SIGV random. * ggml_debug: ci: add tests * ggml_debug: EOL in CMakeLists.txt * ggml_debug: Remove unused param n_batch, no batching here * ggml_debug: fix trailing spaces * ggml_debug: fix trailing spaces * common: fix cb_eval and user data not initialized * ci: build revert label * ggml_debug: add main test label * doc: add a model: add a link to ggml-debug * ggml-debug: add to make toolchain * ggml-debug: tests add the main label * ggml-debug: ci add test curl label * common: allow the warmup to be disabled in llama_init_from_gpt_params * ci: add curl test * ggml-debug: better tensor type support * gitignore : ggml-debug * ggml-debug: printing also the sum of each tensor * ggml-debug: remove block size * eval-callback: renamed from ggml-debug * eval-callback: fix make toolchain --------- Co-authored-by: slaren <slarengh@gmail.com> Co-authored-by: Georgi Gerganov <ggerganov@gmail.com>	2024-04-11 14:51:07 +02:00
Jared Van Bortel	1b67731e18	BERT tokenizer fixes (#6498 ) Key changes: * BERT conversion: fix abuse of LlamaHfVocab, do not set BOS or EOS * Nomic Embed conversion: pad vocab instead of slicing embedding tensor * llama_tokenize: handle added special tokens like HF does	2024-04-09 13:44:08 -04:00
Rick G	e3c337d87c	llama : support negative ith in llama_get_ API (#6519 ) * llama_sampling_sample with default args is more naively usable * Batches populated by either llama_batch_get_one or llama_batch_add work with default args * Previously get_one could use the default argument * Previously add should usually have used the last index where logits[idx] == true * This hopefully encourages the use of llama_batch_add * By giving expected results when using default arguments. * Adds "negative indexing" feature to llama_get_logits_ith and llama_get_embeddings_ith * Believed to work with any currently well behaved program * Default arg now works for both cases (previously would give strange results for add case) * Any non-negative number is unaffected and behaves as previously * Negative arguments were previously invalid. * Implemented as a special case of indexing as suggested by @compilade in https://github.com/ggerganov/llama.cpp/pull/6519 * Fixed mismatch type errors * cited in macOS CI tests * Missed in original updates based on PR feedback in https://github.com/ggerganov/llama.cpp/pull/6519	2024-04-08 16:02:30 +03:00
Jan Boon	beea6e1b16	llama : save and restore kv cache for single seq id (#6341 ) * llama : save and restore kv cache for single seq id * remove trailing whitespace * respond error in case there's no space in the kv cache * add kv seq save restore to test case * add --slot-save-path arg to enable save restore and restrict save location * Returning 0 for some cases, instead of asserting. * cleanup error cases * rename sequence state functions * rename state get set functions * add previous function names back in with DEPRECATED notice * update doc * adjust endpoints to preferred style * fix restoring zero cell count * handle seq rm return value * unused param * keep in the size check * fix return types * add server test case for slot save restore * cleanup * add cake * cleanup style * add special * removing a whole sequence never fails * move sequence state file functionality from server to llama to match session api and add version tags * catch exceptions on save as well * error log messages * check types for stricter restore * update server doc * readme : update API changes date * strict filename validation * move include, reject bom as well * also reject empty filename * reject whitespace and trailing dot --------- Co-authored-by: Martin Evans <martindevans@gmail.com> Co-authored-by: Georgi Gerganov <ggerganov@gmail.com>	2024-04-08 15:43:30 +03:00
Daniel Bevenius	4bcd6b959c	common: remove duplicate check for curl (#6471 ) This commit removes one of the two identical checks for curl being NULL in llama_load_model_from_url. Signed-off-by: Daniel Bevenius <daniel.bevenius@gmail.com>	2024-04-04 09:49:21 +02:00
Sigbjørn Skjæret	e562b9714b	common : change --no-penalize-nl to --penalize-nl (#6334 ) * Change --no-penalize-nl to --penalize-nl * Update documentation too	2024-03-27 09:23:10 +02:00
slaren	280345968d	cuda : rename build flag to LLAMA_CUDA (#6299 )	2024-03-26 01:16:01 +01:00
Neo Zhang Jianyu	95ad616cdd	[SYCL] fix SYCL backend build on windows is break by LOG() error (#6290 ) * fix LOG() error for SYCL, enhance erro check by CI * rollback to bash * add newline at end of file	2024-03-25 15:52:41 +08:00
Minsoo Cheong	64e7b47c69	examples : add "retrieval" (#6193 ) * add `retrieval` example * add README * minor fixes * cast filepos on print * remove use of variable sized array * store similarities in separate vector * print error on insufficient batch size * fix error message printing * assign n_batch value to n_ubatch * fix param definitions * define retrieval-only parameters in retrieval.cpp * fix `--context-file` option to be provided multiple times for multiple files * use vector for `query_emb` * add usage description in README * fix merge conflict * fix usage printing * remove seed setting * fix lint * increase file read buffer size * retrieval : minor --------- Co-authored-by: Georgi Gerganov <ggerganov@gmail.com>	2024-03-25 09:38:22 +02:00
Johannes Gäßler	7aed0ffe68	Fixed lookup compilation issues on Windows (#6273 )	2024-03-24 14:21:17 +01:00
Minsoo Cheong	586e7bc561	sampling : deduplicated code for probability distribution access (#6240 ) * sampling: remove duplicated code for probability distribution access * free original_logits * fix original_logits allocation * fixes based on review @cebtenzzre * change function name to `llama_sampling_prepare`	2024-03-24 10:54:07 +02:00

1 2 3 4 5 ...

312 Commits