ik_llama.cpp

mirror of https://github.com/ikawrakow/ik_llama.cpp.git synced 2026-04-27 18:01:45 +00:00

Files

Kawrakow fc8920282f iqk_mul_mat(ARM_NEON): adding bf16 support (#41 )

It looks like ArmV8 ISA has support for bf16, but my M2 Max
does not have it, so resorting to bf16 -> f32 conversion and
computations in f32. This is 2x slower than f16, but 8x better
compared to what I get if I try to run a bf16 model on the M2
(NEON and Metal).

Co-authored-by: Iwan Kawrakow <iwan.kawrakow@gmail.com>

2024-09-16 16:47:36 +03:00

cmake

Merge mainline llama.cpp (#3 )

2024-07-27 07:55:01 +02:00

include

Adding IQ1_TN - 1.6875 bpw for TriLM ternary models (#44 )

2024-09-09 14:56:34 +03:00

src

iqk_mul_mat(ARM_NEON): adding bf16 support (#41 )

2024-09-16 16:47:36 +03:00

.gitignore

Merge mainline llama.cpp (#3 )

2024-07-27 07:55:01 +02:00

CMakeLists.txt

Merge mainline - Aug 12 2024 (#17 )

2024-08-12 15:14:32 +02:00