ggml : add Q5_0 and Q5_1 quantization (#1187)

* ggml : add Q5_0 quantization (cuBLAS only) * ggml : fix Q5_0 qh -> uint32_t * ggml : fix q5_0 histogram stats * ggml : q5_0 scalar dot product * ggml : q5_0 ARM NEON dot * ggml : q5_0 more efficient ARM NEON using uint64_t masks * ggml : rename Q5_0 -> Q5_1 * ggml : adding Q5_0 mode * quantize : add Q5_0 and Q5_1 to map * ggml : AVX2 optimizations for Q5_0, Q5_1 (#1195) --------- Co-authored-by: Stephan Walter <stephan@walter.name>
author: Georgi Gerganov <ggerganov@gmail.com> 2023-04-26 23:14:13 +0300
committer: GitHub <noreply@github.com> 2023-04-26 23:14:13 +0300
commit: 574406dc7e350ddbffaeca33bf0392b7bfeb1436 (patch)
tree: 03c50ad8b07a612b2169b0bba6b08bd20b11d83a /examples/quantize
parent: 87a6f846d3e929632c45916dd08f1e2a9c72d2a3 (diff)
1 files changed, 2 insertions, 0 deletions
diff --git a/examples/quantize/quantize.cpp b/examples/quantize/quantize.cpp
index ec7f91a..6096659 100644
--- a/examples/quantize/quantize.cpp
+++ b/examples/quantize/quantize.cpp
@@ -10,6 +10,8 @@ static const std::map<std::string, enum llama_ftype> LLAMA_FTYPE_MAP = {
   {"q4_1", LLAMA_FTYPE_MOSTLY_Q4_1},
   {"q4_2", LLAMA_FTYPE_MOSTLY_Q4_2},
   {"q4_3", LLAMA_FTYPE_MOSTLY_Q4_3},
+  {"q5_0", LLAMA_FTYPE_MOSTLY_Q5_0},
+  {"q5_1", LLAMA_FTYPE_MOSTLY_Q5_1},
   {"q8_0", LLAMA_FTYPE_MOSTLY_Q8_0},
 };
author	Georgi Gerganov <ggerganov@gmail.com>	2023-04-26 23:14:13 +0300
committer	GitHub <noreply@github.com>	2023-04-26 23:14:13 +0300
commit	574406dc7e350ddbffaeca33bf0392b7bfeb1436 (patch)
tree	03c50ad8b07a612b2169b0bba6b08bd20b11d83a /examples/quantize
parent	87a6f846d3e929632c45916dd08f1e2a9c72d2a3 (diff)