llama.cpp.git - llama.cpp

Age	Commit message (Collapse)	Author
2023-05-02	Process escape sequences given in prompts (#1173)	DannyDaemonic

2023-05-02	Handle signals properly on Windows (#1123)	DannyDaemonic

2023-05-03	fix missing parameters in `llama_init_from_gpt_params` (#1293)	slaren

2023-05-02	examples : add llama_init_from_gpt_params() common function (#1290)	Ron Evans
	Signed-off-by: deadprogram <ron@hybridgroup.com>
2023-05-02	llama : fix compile warnings	Georgi Gerganov

2023-05-02	examples : improve vertical alignment of a few variables (#1286)	Ron Evans
	Signed-off-by: deadprogram <ron@hybridgroup.com>
2023-05-02	llama : allow 0 as a seed number. (#1275)	Robert Brisita

2023-05-02	main : switch input_noecho to input_echo to remove negation (#979)	Ron Evans
	Signed-off-by: deadprogram <ron@hybridgroup.com>
2023-05-01	Add git-based build information for better issue tracking (#1232)	DannyDaemonic
	* Add git-based build information for better issue tracking * macOS fix * "build (hash)" and "CMAKE_SOURCE_DIR" changes * Redo "CMAKE_CURRENT_SOURCE_DIR" and clearer build messages * Fix conditional dependency on missing target * Broke out build-info.cmake, added find_package fallback, and added build into to all examples, added dependencies to Makefile * 4 space indenting for cmake, attempt to clean up my mess in Makefile * Short hash, less fancy Makefile, and don't modify build-info.h if it wouldn't change it
2023-05-01	llama : fix session load / save (#1263)	Georgi Gerganov

2023-04-30	common : better default number of threads (#934)	jon-chuang
	* commit * fix * try-catch * apply code review * improve * improve * add macos headers * done * remove color * fix windows * minor * fix * Apply suggestions from code review Co-authored-by: DannyDaemonic <DannyDaemonic@gmail.com> * remove * minor * minor --------- Co-authored-by: jon-chuang <jon-chuang@users.noreply.github.com> Co-authored-by: DannyDaemonic <DannyDaemonic@gmail.com>
2023-04-30	Various fixes to mat_mul benchmark (#1253)	Stephan Walter

2023-04-29	build : fix reference to old llama_util.h	Georgi Gerganov

2023-04-29	examples : fix save-load-state + rename llama-util.h	Georgi Gerganov

2023-04-29	common : change default parameters to pre-#1126 (#1223)	Georgi Gerganov

2023-04-29	llama : new sampling algorithms (#1126)	Ivan Stepanov
	* Sample interface, new samplers. New samplers: - locally typical sampling - tail free sampling - frequency and presence penalty - mirostat Ignore EOS fix: -inf should be used. * mirostat * Added --logit-bias and --no-penalize-nl, removed std::span * Use C++11, clarify llama API documentation, rename Mirostat parameters to --mirostat_lr and --mirostat_ent, add temperature sampling for Mirostat, simplify Mirostat sampling API parameters (removed N and k) Use C++11, clarify llama API documentation, rename Mirostat parameters to --mirostat_lr and --mirostat_ent, add temperature sampling for Mirostat, simplify Mirostat sampling API parameters (removed N and k) * Save and load example adjust * Tests * Windows build fix * Windows test fix
2023-04-28	Remove Q4_3 which is no better than Q5 (#1218)	Stephan Walter

2023-04-28	examples : add Jeopardy example (#1168)	CRD716
	* Basic Setup * Prevent Results.txt from coming up * Prefixes, Line separators, etc * editorcheck * introduction to give more consistent results * Basic graph thing * Grading, ready for testing! * Y'all ready to get funky? * fix column removal stuff * missed a few
2023-04-28	llama : add session file format and saved sessions in main (#1169)	Evan Jones

2023-04-26	ggml : add Q5_0 and Q5_1 quantization (#1187)	Georgi Gerganov
	* ggml : add Q5_0 quantization (cuBLAS only) * ggml : fix Q5_0 qh -> uint32_t * ggml : fix q5_0 histogram stats * ggml : q5_0 scalar dot product * ggml : q5_0 ARM NEON dot * ggml : q5_0 more efficient ARM NEON using uint64_t masks * ggml : rename Q5_0 -> Q5_1 * ggml : adding Q5_0 mode * quantize : add Q5_0 and Q5_1 to map * ggml : AVX2 optimizations for Q5_0, Q5_1 (#1195) --------- Co-authored-by: Stephan Walter <stephan@walter.name>
2023-04-26	quantize : use `map` to assign quantization type from `string` (#1191)	Pavol Rusnak
	instead of `int` (while `int` option still being supported) This allows the following usage: `./quantize ggml-model-f16.bin ggml-model-q4_0.bin q4_0` instead of: `./quantize ggml-model-f16.bin ggml-model-q4_0.bin 2`
2023-04-25	ggml : add Q8_0 quantization format (rename the old one to Q8_1) (ARM NEON) ↵	Georgi Gerganov
	(#1179) * ggml : add Q8_0 quantization format (rename the old one to Q8_1) * tests : fix test-quantize-fns * ggml : finalize Q8_0 implementation * ggml : use q4_0_q8_0 and q4_2_q8_0 * ggml : fix Q8_0 dot product bug (ARM) * ggml : Q8_0 unroll x2 * ggml : fix bug - using wrong block type * ggml : extend quantize_fns_t with "vec_dot_type" * ggml : fix Q8_0 to use 255 values out of 256 * ggml : fix assert using wrong QK4_2 instead of QK4_3
2023-04-24	examples : add save_load_state example (#1150)	xaedes
	* add save_load_state example * use <cstdio> instead of <iostream> and fprintf / printf instead of cout * renamed save-load-state example files replacing underscores by dashes
2023-04-24	examples/main README improvements and some light refactoring (#1131)	mgroeber9110

2023-04-23	Fix LoRA acronym (#1145)	slaren

2023-04-23	Added README.md for main with examples and explanations (#1139)	DannyDaemonic

2023-04-22	Fix CI: ARM NEON, quantization unit tests, editorconfig (#1122)	Stephan Walter

2023-04-22	llama : print timings on ctrl+c exit (#1021)	wbpxre150
	* print timings on ctrl+c exit * remove redundant free memory call. * add global pointer to ctx.
2023-04-22	llama : have n_batch default to 512 (#1091)	eiery
	* set default n_batch to 512 when using BLAS * spacing * alternate implementation of setting different n_batch for BLAS * set n_batch to 512 for all cases
2023-04-22	examples : Improve Alpaca Default Repeat Penalty: Better Match Alpaca.cpp ↵	Clint Herron
	Experience (#1107) * Moving parameters to separate lines for readability. * Increasing repeate_penalty to 1.1 to make alpaca more usable by default. * Adding trailing newline.
2023-04-21	main : evaluate tokens in batches after swapping context (#1014)	Alex Klinkhamer
	* examples : evaluate tokens in batches after swapping context * Update examples/main/main.cpp --------- Co-authored-by: Georgi Gerganov <ggerganov@gmail.com>
2023-04-21	Show perplexity ETA in hours and minutes (#1096)	slaren

2023-04-20	llama : multi-threaded quantization (#1075)	Kawrakow
	* Multi-threading quantization. Not much gain for simple quantizations, bit it will be important for quantizations that require more CPU cycles. * Multi-threading for quantize-stats It now does the job in ~14 seconds on my Mac for Q4_0, Q4_1 and Q4_2. Single-threaded it was taking more than 2 minutes after adding the more elaborate version of Q4_2. * Reviewer comments * Avoiding compiler confusion After changing chunk_size to const int as suggested by @ggerganov, clang and GCC starting to warn me that I don't need to capture it in the lambda. So, I removed it from the capture list. But that makes the MSVC build fail. So, making it a constexpr to make every compiler happy. * Still fighting with lambda captures in MSVC --------- Co-authored-by: Iwan Kawrakow <iwan.kawrakow@gmail.com> Co-authored-by: Georgi Gerganov <ggerganov@gmail.com>
2023-04-20	ggml : add Q4_3 quantization (#1082)	Georgi Gerganov

2023-04-18	ggml : add new Q4_2 quantization (ARM only) (#1046)	Georgi Gerganov
	* ggml : Q4_2 ARM * ggml : add ggml_is_quantized() * llama : update llama_type_name() with Q4_2 entry * ggml : speed-up q4_2 - 4 threads: ~100ms -> ~90ms - 8 threads: ~55ms -> ~50ms * ggml : optimize q4_2 using vmlaq_n_f32 + vmulq_n_f32
2023-04-17	Add LoRA support (#820)	slaren

2023-04-17	quantize-stats : fix bug in --type argument	Georgi Gerganov

2023-04-16	examples: add missing <ctime> include for time() (#1011)	Pavol Rusnak

2023-04-15	benchmark : fix result validation in benchmark-q4_0-matmult (#987)	Ivan Komarov

2023-04-14	Revert "main : alternative instruct mode (Vicuna support, etc.) (#863)" (#982)	Pavol Rusnak
	This reverts commit f4d277ae17247ee51129ef1a9ff74d377cc90b1b.
2023-04-14	Expose type name from ggml (#970)	Pavol Rusnak
	Avoid duplication of type names in utils Co-authored-by: Håkon H. Hitland <haakon@likedan.net>
2023-04-14	main : alternative instruct mode (Vicuna support, etc.) (#863)	Tomáš Pazdiora
	* Add support for configs, add configurable prefixes / suffixes, deprecate instruct mode, add stop prompt * Add multiline mode, update text input. * bugfix * update implementation * typos * Change --multiline implementation to be toggled by EOF. * bugfix * default multiline mode * add more configs * update formating * update formatting * apply suggestions
2023-04-14	perplexity : add support for batch size to `--perplexity` (#407)	Gary Linscott
	* Add support to batch size for perplexity * Revert "Fix memory allocation issues and seg faults" This reverts commit 4870e455b3653f7d7769fa5772b2c90ffad088df. * update from merge * Remove perplexity from main * updates * Update batch size for efficiency
2023-04-13	common : remove unnecessary includes (#947)	CRD716

2023-04-13	llama : merge llama_internal.h into llama.h	Georgi Gerganov
	Hide it behind an #ifdef
2023-04-13	fix whitespace (#944)	CRD716

2023-04-13	examples : add -n to alpaca and gpt4all scripts (#706)	niansa/tuxifan

2023-04-13	benchmark : add tool for timing q4_0 matrix multiplication (#653)	SebastianApel
	* Initial version of q4_0 matrix multiplication benchmark * Bugfix: Added dependency to ggml.o to benchmark * Reviewer requests: added parameter for threads, switched to ggml_time_us() * Reviewer input: removed rtsc, use epsilon for check * Review comment: Removed set_locale * Feature: Param for numer of iterations, Bugfix for use of parameter threads * Reviewer suggestion: Moved to examples * Reviewer feedback: Updated clean: and benchmark: sections --------- Co-authored-by: Georgi Gerganov <ggerganov@gmail.com>
2023-04-11	Fix whitespace, add .editorconfig, add GitHub workflow (#883)	Pavol Rusnak

2023-04-11	Add enum llama_ftype, sync ggml_type to model files (#709)	Stephan Walter