llama.cpp.git - llama.cpp

Age	Commit message (Collapse)	Author
2023-07-06	convert : update for baichuan (#2081)	Judd
	1. guess n_layers; 2. relax warnings on context size; 3. add a note that its derivations are also supported. Co-authored-by: Judd <foldl@boxvest.com>
2023-06-29	Use unsigned for random seed (#2006)	Howard Su
	* Use unsigned for random seed. Keep -1 as the value to use a time based seed. Co-authored-by: Georgi Gerganov <ggerganov@gmail.com>
2023-06-26	ggml : add NUMA support (#1556)	zrm
	* detect NUMA systems and pin work threads to nodes (linux) * disable mmap prefetch/readahead for NUMA systems * avoid sending finalize op to thread pool if it does nothing * silence robot * fix args * make --numa a param * recommendation that n_nodes evenly divide n_threads did not warrant such aggressive enforcement * lower synchronization overhead * statically allocate * move numa state to g_state * add description for --numa * ggml : minor style changes * ggml : minor style + try fix sanitizer build * llama : allow to initialize backend with NUMA support * llama : avoid ggml include in llama-util.h * ggml : style / formatting * ggml : fix handling of ops with n_threads > n_tasks > 1 * server : utilize numa parameter --------- Co-authored-by: Georgi Gerganov <ggerganov@gmail.com>
2023-06-24	llama : make model stateless and context stateful (llama_state) (#1797)	Didzis Gosko
	* llama : make model stateless and context stateful * llama : minor cleanup * llama : update internal API declaration * Apply suggestions from code review fix style Co-authored-by: Georgi Gerganov <ggerganov@gmail.com> * Missing model memory release * Fix style * Add deprecated warning for public API function llama_init_from_file * Update public API use cases: move away from deprecated llama_init_from_file * Deprecate public API function llama_apply_lora_from_file --------- Co-authored-by: Georgi Gerganov <ggerganov@gmail.com>
2023-06-17	minor : warning fixes	Georgi Gerganov

2023-06-16	Fixed possible macro redefinition (#1892)	FrankHB
	MinGW libstdc++ may define `NOMINMAX` unconditionally. This fixes the case when it is already defined.
2023-06-16	build : fix and ignore MSVC warnings (#1889)	Borislav Stanimirov

2023-06-14	CUDA full GPU acceleration, KV cache in VRAM (#1827)	Johannes Gäßler
	* Fixed CUDA RoPE * ggml_cuda_mul_mat_vec_p021 * ggml_cuda_scale * ggml_cuda_diag_mask_inf * ggml_is_permuted * ggml_cuda_cpy * flatten rows for ggml_cuda_op * Added a --low-vram option * Fixed Windows performance * Fixed LLAMA_CUDA_DMMV_Y > 1 for WizardLM
2023-06-13	llama : do a warm-up eval at start for better timings (#1824)	Georgi Gerganov

2023-06-11	Fix issue where interactive mode crashes when input exceeds ctx size (#1789)	Kerfuffle
	* Fix issue where interactive mode in the main example crashes when input exceeds ctx size * Ensure the context size is at least 8 tokens in the main example. Closes #1768
2023-06-06	main: add the possibility to open the prompt cache read-only (#1640)	Willy Tarreau
	The prompt cache constitutes a nice speed up when using the same prompt prefix across multiple evaluations, but when using it, it will also be updated, which is not always desirable. One use case is to have a large prompt containing some context and usage rules, and a second part containing variable data of the problem being studied. In this case it's desirable to be able to save the first part once, and to always reuse it as-is without updating it with the second part. The new argument --prompt-cache-ro enables this read-only mode on the prompt cache. The prompt's contents that match the cache are loaded from the cache but the rest is not modified. This allowed to reduce a total analysis time from 112s to 49.7s here, without having to backup and restore a copy of the prompt, which takes significant time at 500 MB. Signed-off-by: Willy Tarreau <w@1wt.eu>
2023-06-06	Multi GPU support, CUDA refactor, CUDA scratch buffer (#1703)	Johannes Gäßler
	* CUDA multi GPU + scratch ggml_cuda_compute_forward Tensor parallelism ggml_cuda_add ggml_cuda_rms_norm ggml_cuda_silu CUDA scratch buffer --main-gpu CLI option
2023-06-04	llama : Metal inference (#1642)	Georgi Gerganov
	* mtl : export the LLaMA computation graph * ci : disable temporary * mtl : adapt the MNIST example as starter * mtl : no need for mtl-export tool, add cli arg for main instead * mtl : export just a small part of the graph for now to make it easier * mtl : move MSL code into separate file for easy editing * mtl : initial get_rows_q4_0 kernel * mtl : confirmed get_rows_q4_0 is working correctly * mtl : add rms_norm kernel + confirm working * mtl : add mul kernel + confirm working * mtl : initial mul_mat Q4 kernel (wrong results) * mtl : mul_mat fixes (still wrong) * mtl : another mul_mat Q4 (still does not work) * mtl : working mul_mat q4 * ggml : fix handling of "view" ops in ggml_graph_import() * mtl : add rope kernel * mtl : add reshape and transpose handling * ggml : store offset as opt arg for ggml_view_xd() operators * mtl : add cpy kernel + handle view ops * mtl : confirm f16 x f32 attention mul mat * mtl : add scale kernel * mtl : add diag_mask_inf kernel * mtl : fix soft_max kernel * ggml : update ggml_nbytes() to handle non-contiguous tensors * mtl : verify V tensor contents * mtl : add f32 -> f32 cpy kernel * mtl : add silu kernel * mtl : add non-broadcast mul kernel * mtl : full GPU inference of the computation graph * mtl : optimize rms_norm and soft_max kernels * mtl : add f16 mat x f32 vec multiplication kernel * mtl : fix bug in f16 x f32 mul mat + speed-up computation * mtl : faster mul_mat_q4_0_f32 kernel * mtl : fix kernel signature + roll inner loop * mtl : more threads for rms_norm + better timing * mtl : remove printfs from inner loop * mtl : simplify implementation * mtl : add save/load vocab to ggml file * mtl : plug Metal inference into llama.cpp (very quick-n-dirty) * mtl : make it work with main example Lots of hacks but at least now it generates text * mtl : preparing for merge * mtl : clean-up ggml mtl interface + suport scratch / inplace * mtl : remove temp / debug code * metal : final refactoring and simplification * Revert "ci : disable temporary" This reverts commit 98c267fc77fe811082f672538fc91bcfc9072d63. * metal : add comments * metal : clean-up stuff, fix typos * readme : add Metal instructions * readme : add example for main
2023-06-03	Fix prompt cache saving and chat-persistent rollover (#1678)	Evan Jones
	* Fix prompt cache saving and chat-persistent rollover (fixes #1670) * clang-tidy Co-authored-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com> --------- Co-authored-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
2023-05-29	Work around for recalculating logits in cached prompts (Fixes #1585) (#1609)	DannyDaemonic
	* Work around for recalculating logits in cached prompts
2023-05-28	Only show -ngl option when relevant + other doc/arg handling updates (#1625)	Kerfuffle
	1. Add a `LLAMA_SUPPORTS_GPU_OFFLOAD` define to `llama.h` (defined when compiled with CLBlast or cuBLAS) 2. Update the argument handling in the common example code to only show the `-ngl`, `--n-gpu-layers` option when GPU offload is possible. 3. Add an entry for the `-ngl`, `--n-gpu-layers` option to the `main` and `server` examples documentation 4. Update `main` and `server` examples documentation to use the new style dash separator argument format 5. Update the `server` example to use dash separators for its arguments and adds `-ngl` to `--help` (only shown when compiled with appropriate support). It will still support `--memory_f32` and `--ctx_size` for compatibility. 6. Add a warning discouraging use of `--memory-f32` for the `main` and `server` examples `--help` text as well as documentation. Rationale: https://github.com/ggerganov/llama.cpp/discussions/1593#discussioncomment-6004356
2023-05-25	Some improvements to loading the session with --prompt-cache (#1550)	Kerfuffle
	Improvements to loading the session with `--prompt-cache` in the `main` example. 1. Fix an issue where the `--seed` parameter was ignored when loading a cached prompt. 2. When loading a cached prompt, you previously had to specify the saved prompt (or a prefix of it) again. This pull changes that behavior to default to the prompt that was cached if a prompt wasn't specified by the user.
2023-05-20	llama : add llama_init_backend() API (close #1527)	Georgi Gerganov

2023-05-19	main : make reverse prompt option act as a stop token in non-interactive ↵	Jason McCartney
	mode (#1032) * Make reverse prompt option act as a stop token in non-interactive scenarios * Making requested review changes * Update gpt_params_parse and fix a merge error * Revert "Update gpt_params_parse and fix a merge error" This reverts commit 2bb2ff1748513591ad45b175a75ed1d8089d84c8. * Update gpt_params_parse and fix a merge error take 2
2023-05-18	Fixes #1511 lambda issue for w64devkit (mingw) (#1513)	DannyDaemonic
	* Fix for w64devkit and mingw
2023-05-16	define default model path once, sync path with readme (#1366)	András Salamon

2023-05-12	llama : fix --mtest option (close #1414)	Georgi Gerganov

2023-05-10	main : add option to save full output to session (#1338)	Evan Jones
	* main : add option to save full output to session * split behavior into --session and --prompt-cache * restore original implementation with new names * PR comments * move the check for incompatible parameters to gpt_params_parse * Fix whitespace Co-authored-by: DannyDaemonic <DannyDaemonic@gmail.com> --------- Co-authored-by: DannyDaemonic <DannyDaemonic@gmail.com>
2023-05-08	Interface improvements and `--multiline-input` (previously `--author-mode`) ↵	DannyDaemonic
	(#1040) * Interface improvements * Multiline input * Track character width * Works with all characters and control codes + Windows console fixes
2023-05-08	llama : require first token to be BOS (#1303)	Georgi Gerganov
	* llama : require first token to be BOS * scripts : add ppl-run-all.sh * perplexity : add BOS for each chunk * readme : update perplexity values after BOS fix * perplexity : add clarifying comments
2023-05-06	Remove default arguments from sampling functions (#1343)	Jed Fox

2023-05-04	main : add --in-suffix option (#1318)	44670
	* adding --in-suffix option * print input suffix before generation
2023-05-04	Only escape prompts when used with `-e` (#1311)	DannyDaemonic

2023-05-04	Update main's README.md with new features (#1296)	DannyDaemonic

2023-05-04	fix #1224 reverse prompt and multi line (#1297)	Tomas
	* fix reverse prompt and multi line * Code Formatting Co-authored-by: Georgi Gerganov <ggerganov@gmail.com> --------- Co-authored-by: Georgi Gerganov <ggerganov@gmail.com>
2023-05-02	Handle signals properly on Windows (#1123)	DannyDaemonic

2023-05-02	examples : add llama_init_from_gpt_params() common function (#1290)	Ron Evans
	Signed-off-by: deadprogram <ron@hybridgroup.com>
2023-05-02	examples : improve vertical alignment of a few variables (#1286)	Ron Evans
	Signed-off-by: deadprogram <ron@hybridgroup.com>
2023-05-02	llama : allow 0 as a seed number. (#1275)	Robert Brisita

2023-05-02	main : switch input_noecho to input_echo to remove negation (#979)	Ron Evans
	Signed-off-by: deadprogram <ron@hybridgroup.com>
2023-05-01	Add git-based build information for better issue tracking (#1232)	DannyDaemonic
	* Add git-based build information for better issue tracking * macOS fix * "build (hash)" and "CMAKE_SOURCE_DIR" changes * Redo "CMAKE_CURRENT_SOURCE_DIR" and clearer build messages * Fix conditional dependency on missing target * Broke out build-info.cmake, added find_package fallback, and added build into to all examples, added dependencies to Makefile * 4 space indenting for cmake, attempt to clean up my mess in Makefile * Short hash, less fancy Makefile, and don't modify build-info.h if it wouldn't change it
2023-05-01	llama : fix session load / save (#1263)	Georgi Gerganov

2023-04-29	common : change default parameters to pre-#1126 (#1223)	Georgi Gerganov

2023-04-29	llama : new sampling algorithms (#1126)	Ivan Stepanov
	* Sample interface, new samplers. New samplers: - locally typical sampling - tail free sampling - frequency and presence penalty - mirostat Ignore EOS fix: -inf should be used. * mirostat * Added --logit-bias and --no-penalize-nl, removed std::span * Use C++11, clarify llama API documentation, rename Mirostat parameters to --mirostat_lr and --mirostat_ent, add temperature sampling for Mirostat, simplify Mirostat sampling API parameters (removed N and k) Use C++11, clarify llama API documentation, rename Mirostat parameters to --mirostat_lr and --mirostat_ent, add temperature sampling for Mirostat, simplify Mirostat sampling API parameters (removed N and k) * Save and load example adjust * Tests * Windows build fix * Windows test fix
2023-04-28	llama : add session file format and saved sessions in main (#1169)	Evan Jones

2023-04-24	examples/main README improvements and some light refactoring (#1131)	mgroeber9110

2023-04-23	Fix LoRA acronym (#1145)	slaren

2023-04-23	Added README.md for main with examples and explanations (#1139)	DannyDaemonic

2023-04-22	Fix CI: ARM NEON, quantization unit tests, editorconfig (#1122)	Stephan Walter

2023-04-22	llama : print timings on ctrl+c exit (#1021)	wbpxre150
	* print timings on ctrl+c exit * remove redundant free memory call. * add global pointer to ctx.
2023-04-21	main : evaluate tokens in batches after swapping context (#1014)	Alex Klinkhamer
	* examples : evaluate tokens in batches after swapping context * Update examples/main/main.cpp --------- Co-authored-by: Georgi Gerganov <ggerganov@gmail.com>
2023-04-17	Add LoRA support (#820)	slaren

2023-04-16	examples: add missing <ctime> include for time() (#1011)	Pavol Rusnak

2023-04-14	Revert "main : alternative instruct mode (Vicuna support, etc.) (#863)" (#982)	Pavol Rusnak
	This reverts commit f4d277ae17247ee51129ef1a9ff74d377cc90b1b.
2023-04-14	main : alternative instruct mode (Vicuna support, etc.) (#863)	Tomáš Pazdiora
	* Add support for configs, add configurable prefixes / suffixes, deprecate instruct mode, add stop prompt * Add multiline mode, update text input. * bugfix * update implementation * typos * Change --multiline implementation to be toggled by EOF. * bugfix * default multiline mode * add more configs * update formating * update formatting * apply suggestions