ik_llama.cpp

mirror of https://github.com/ikawrakow/ik_llama.cpp.git synced 2026-02-25 07:34:10 +00:00

Files

Iwan Kawrakow bdc882bfac Moving 4D gemm logic from ggml.c to iqk_mul_mat.cpp

This allows us to optimize TG performance for GQA models.
E.g., for IQ4_XS L3-8B with 8k TG-64 goes from 8.6 to 10.26 t/s.

2025-02-14 18:03:19 +02:00

2024-07-27 07:55:01 +02:00

2025-02-09 09:14:52 +02:00

2025-02-14 18:03:19 +02:00

.gitignore

2024-07-27 07:55:01 +02:00

CMakeLists.txt

2025-02-09 18:59:33 +02:00