Files
ik_llama.cpp/examples/quantize/quantize.cpp
Iwan Kawrakow c43e747c7c Adding iq4_xs_r4
This is a 1st working version on Zen4.
We get PP-512(LLaMA-3.1-8B) = 226 t/s, so 16% slower
than iq4_nl_x4.
2024-12-04 09:16:45 +02:00

24 KiB