ik_llama.cpp/ggml/include/ggml.h at 692dc0d9b57e044038d6079e2ed093d0484319b2

mirror of https://github.com/ikawrakow/ik_llama.cpp.git synced 2026-04-29 02:41:47 +00:00

Files

Kawrakow 3248a35992 Adding IQ3_KS quants (#566 )

* iq3_ks: basics

* iq3_ks: CUDA dequantize

* iq3_ks: CUDA mmvq

* iq3_ks: mmq

* iq3_ks: faster mmq

* iq3_ks: Zen4

* iq3_ks: AVX2 convert to q8_k_r8

This gives usPP-512 = 360 t/s.

* iq3_ks: AVX2 GEMM/GEMV

* iq3_ks: NEON GEMM/GEMV

* iq3_ks: NEON convert to q8_k_r8

This gives us PP-512 = 164 t/s.

* iq3_ks: Metal dequantize

* iq3_ks: Metal gemv - pathetic performance

---------

Co-authored-by: Iwan Kawrakow <iwan.kawrakow@gmail.com>

2025-07-02 09:27:47 +02:00

99 KiB

Raw Blame History

View Raw

99 KiB Raw Blame History

99 KiB

Raw Blame History