llama.cpp.git - llama.cpp

Age	Commit message (Collapse)	Author
2023-05-01	cuBLAS: refactor and optimize f16 mat mul performance (#1259)	slaren
	* cuBLAS: refactor, convert fp16 to fp32 on device * cuBLAS: use multiple streams, choose smartly between mul_mat_q and mul_mat_f16 * fix build * cuBLAS: update block_q5_1
2023-05-01	cuBLAS: fall back to pageable memory if pinned alloc fails (#1233)	slaren
	* cuBLAS: fall back to pageable memory if pinned alloc fails * cuBLAS: do not use pinned memory if env variable GGML_CUDA_NO_PINNED is set
2023-04-29	cuBLAS: use host pinned memory and dequantize while copying (#1207)	slaren
	* cuBLAS: dequantize simultaneously while copying memory * cuBLAS: use host pinned memory * cuBLAS: improve ggml_compute_forward_mul_mat_f16_f32 with pinned memory * cuBLAS: also pin kv cache * fix rebase
2023-04-29	cuBLAS: non-contiguous tensor support (#1215)	Henri Vasserman
	* Cuda: non-contiguous tensor support * remove extra stuff * rename * fix error * more fixes, now OpenBLAS and CLBlast build too * now then?
2023-04-28	Remove Q4_3 which is no better than Q5 (#1218)	Stephan Walter

2023-04-26	ggml : add Q5_0 and Q5_1 quantization (#1187)	Georgi Gerganov
	* ggml : add Q5_0 quantization (cuBLAS only) * ggml : fix Q5_0 qh -> uint32_t * ggml : fix q5_0 histogram stats * ggml : q5_0 scalar dot product * ggml : q5_0 ARM NEON dot * ggml : q5_0 more efficient ARM NEON using uint64_t masks * ggml : rename Q5_0 -> Q5_1 * ggml : adding Q5_0 mode * quantize : add Q5_0 and Q5_1 to map * ggml : AVX2 optimizations for Q5_0, Q5_1 (#1195) --------- Co-authored-by: Stephan Walter <stephan@walter.name>
2023-04-25	ggml : add Q8_0 quantization format (rename the old one to Q8_1) (ARM NEON) ↵	Georgi Gerganov
	(#1179) * ggml : add Q8_0 quantization format (rename the old one to Q8_1) * tests : fix test-quantize-fns * ggml : finalize Q8_0 implementation * ggml : use q4_0_q8_0 and q4_2_q8_0 * ggml : fix Q8_0 dot product bug (ARM) * ggml : Q8_0 unroll x2 * ggml : fix bug - using wrong block type * ggml : extend quantize_fns_t with "vec_dot_type" * ggml : fix Q8_0 to use 255 values out of 256 * ggml : fix assert using wrong QK4_2 instead of QK4_3
2023-04-21	Improve cuBLAS performance by using a memory pool (#1094)	slaren
	* Improve cuBLAS performance by using a memory pool * Move cuda specific definitions to ggml-cuda.h/cu * Add CXX flags to nvcc * Change memory pool synchronization mechanism to a spin lock General code cleanup
2023-04-20	Add Q4_3 support to cuBLAS (#1086)	slaren

2023-04-20	Improve cuBLAS performance by dequantizing on the GPU (#1065)	slaren