Adding q5_0_r4

We get PP-512(LLaMA-3.1-8B) = 256.7 t/s on a Ryzen-7950X. We even get TG-128 improvement to 11.7 t/s from 11.1 t/s.
2026-02-27 16:44:21 +00:00 · 2024-12-03 11:29:57 +02:00
parent ccec00939a
commit fad847d753
10 changed files with 323 additions and 21 deletions
--- a/include/llama.h
+++ b/include/llama.h
@@ -182,6 +182,7 @@ extern "C" {
                                                //
        LLAMA_FTYPE_MOSTLY_Q4_0_R4       = 202, // except 1d tensors
        LLAMA_FTYPE_MOSTLY_Q8_0_R4       = 207, // except 1d tensors
+        LLAMA_FTYPE_MOSTLY_Q5_0_R4       = 208, // except 1d tensors
        LLAMA_FTYPE_MOSTLY_IQ4_NL_X4     = 225, // except 1d tensors

        LLAMA_FTYPE_GUESSED = 1024, // not specified in the model file