ik_llama.cpp

mirror of https://github.com/ikawrakow/ik_llama.cpp.git synced 2026-02-25 15:44:10 +00:00

Files

Iwan Kawrakow 5de1cf4885 Faster iq4_xs_r4 on Zen4

The trick is to simply prepare the Q8 block sums for
blocks of 32 as floats. This brings PP-512 up to 254.6 t/s
from 224 t/s.

2024-12-08 15:44:49 +02:00

ggml-alloc.h

2024-07-27 07:55:01 +02:00

ggml-backend.h

Bitnet changes (#106 )

2024-10-25 13:08:43 +02:00

ggml-blas.h

2024-07-27 07:55:01 +02:00

ggml-cann.h

2024-07-27 07:55:01 +02:00

ggml-cuda.h

2024-08-12 15:14:32 +02:00

ggml-kompute.h

2024-07-27 07:55:01 +02:00

ggml-metal.h

2024-08-12 15:14:32 +02:00

ggml-rpc.h

2024-07-27 07:55:01 +02:00

ggml-sycl.h

2024-07-27 07:55:01 +02:00

ggml-vulkan.h

2024-07-27 07:55:01 +02:00

ggml.h

Faster iq4_xs_r4 on Zen4

2024-12-08 15:44:49 +02:00