ik_llama.cpp/examples/quantize/quantize.cpp at 525ad3b4c1d50ef4c0c6f6ba7ad8dabdd62f3d37

mirror of https://github.com/ikawrakow/ik_llama.cpp.git synced 2026-03-04 02:50:01 +00:00

Files

Iwan Kawrakow c43e747c7c Adding iq4_xs_r4

This is a 1st working version on Zen4.
We get PP-512(LLaMA-3.1-8B) = 226 t/s, so 16% slower
than iq4_nl_x4.

2024-12-04 09:16:45 +02:00

View Raw