diff options
author | Johannes Gäßler <johannesg@5d6.de> | 2023-06-14 19:47:19 +0200 |
---|---|---|
committer | GitHub <noreply@github.com> | 2023-06-14 19:47:19 +0200 |
commit | 254a7a7a5ff4c874ff8488f1f5cbdd7e9c89d682 (patch) | |
tree | 65f35a2d189f3cf6f1f625b2acb343c2dd77790d /docs | |
parent | 92549202659fc23ba9fec5e688227d0da9b06b40 (diff) |
CUDA full GPU acceleration, KV cache in VRAM (#1827)
* Fixed CUDA RoPE
* ggml_cuda_mul_mat_vec_p021
* ggml_cuda_scale
* ggml_cuda_diag_mask_inf
* ggml_is_permuted
* ggml_cuda_cpy
* flatten rows for ggml_cuda_op
* Added a --low-vram option
* Fixed Windows performance
* Fixed LLAMA_CUDA_DMMV_Y > 1 for WizardLM
Diffstat (limited to 'docs')
0 files changed, 0 insertions, 0 deletions