llama.cpp.git - llama.cpp

Age	Commit message (Collapse)	Author
2023-04-08	Add quantize-stats command for testing quantization (#728)	unbounded
	Command that calculates some statistics over the errors introduced by quantization, like mean square error, max error and some percentile errors for layer weights. Should be useful for testing quantization improvements. Exposes some internal state from ggml and llama for testing
2023-04-07	make : add libllama.so target for llama-cpp-python (#797)	bhubbb
	I was able to get llama-cpp-python working but only when I build libllama.so with make.
2023-04-05	make : missing host optimizations in CXXFLAGS (#763)	Ivan Stepanov

2023-04-02	make : use -march=native -mtune=native on x86 (#609)	Fabian

2023-03-30	make : fix darwin f16c flags check (#615)	david raistrick
	...there was no check. ported upstream from https://github.com/zanussbaum/gpt4all.cpp/pull/2 (I dont see any clean path for upstream patches)
2023-03-28	all : be more strict about converting float to double (#458)	Stephan Walter
	* Be more strict about converting float to double * Test equivalence of round, SILU implementations Test module is commented out in CMakeLists.txt because the tests may take a long time, depending on how much the compiler optimizes. * Fix softmax in perplexity.cpp * all : prefer float over double where appropriate * perplexity : add <cmath> --------- Co-authored-by: Georgi Gerganov <ggerganov@gmail.com>
2023-03-28	Add embedding example to Makefile (#540)	RJ Adriaansen

2023-03-25	Overhaul the examples structure	Georgi Gerganov
	- main -> examples - utils -> examples (renamed to "common") - quantize -> examples - separate tools for "perplexity" and "embedding" Hope I didn't break something !
2023-03-24	additional optimizations for POWER9 (#454)	Cameron Kaiser

2023-03-23	Fix Makefile echo escape codes (by removing them). (#418)	Kerfuffle

2023-03-22	Introduce C-style API (#370)	Georgi Gerganov
	* Major refactoring - introduce C-style API * Clean up * Add <cassert> * Add <iterator> * Add <algorithm> .... * Fix timing reporting and accumulation * Measure eval time only for single-token calls * Change llama_tokenize return meaning
2023-03-21	makefile: Fix CPU feature detection on Haiku (#218)	Alex von Gluck IV

2023-03-21	Add OpenBSD support (#314)	Kevin Lo

2023-03-21	Makefile: slightly cleanup for Mac Intel; echo instead of run ./main -h (#335)	Qingyou Meng

2023-03-21	Add tokenizer test + revert to C++11 (#355)	Georgi Gerganov
	* Add test-tokenizer-0 to do a few tokenizations - feel free to expand * Added option to convert-pth-to-ggml.py script to dump just the vocabulary * Added ./models/ggml-vocab.bin containing just LLaMA vocab data (used for tests) * Added utility to load vocabulary file from previous point (temporary implementation) * Avoid using std::string_view and drop back to C++11 (hope I didn't break something) * Rename gpt_vocab -> llama_vocab * All CMake binaries go into ./bin/ now
2023-03-21	Add initial AVX512 support for dot product on Linux (#320)	Casey Primozic
	* Update Makefile to detect AVX512 support and add compiler flags if it's available * Based on existing AVX2 implementation, dot product on one 32-value block of 4-bit quantized ints at a time * Perform 8 bit -> 16 bit sign extension and multiply+add on 32 values at time instead of 16 * Use built-in AVX512 horizontal reduce add to get sum at the end * Manual unrolling on inner dot product loop to reduce loop counter overhead
2023-03-20	sentencepiece bpe compatible tokenizer (#252)	Mack Straight
	* potential out of bounds read * fix quantize * style * Update convert-pth-to-ggml.py * mild cleanup * don't need the space-prefixing here rn since main.cpp already does it * new file magic + version header field * readme notice * missing newlines Co-authored-by: slaren <2141330+slaren@users.noreply.github.com>
2023-03-13	Add NetBSD support. (#90)	Thomas Klausner

2023-03-11	Update Makefile var + add comment	Georgi Gerganov

2023-03-10	Initial release	Georgi Gerganov