ik_llama.cpp

mirror of https://github.com/ikawrakow/ik_llama.cpp.git synced 2026-05-11 16:40:16 +00:00

Files

Kawrakow fcd1e124e0 Faster MoE token generation on CUDA (#248 )

* This gives us ~20% TG speedup for DeepSeek on CUDA

* Slightly better

* Also do it for plain (not fused) mul_mat_id

* Guard against numerical precision issues for MLA on CUDA

---------

Co-authored-by: Iwan Kawrakow <iwan.kawrakow@gmail.com>

2025-03-10 16:16:51 +02:00

CMakeLists.txt

Be able to repack tensors at run time (#147 )

2024-12-17 14:16:34 +01:00

llama-grammar.cpp

Merge mainline - Aug 12 2024 (#17 )

2024-08-12 15:14:32 +02:00

llama-grammar.h

Merge mainline llama.cpp (#3 )

2024-07-27 07:55:01 +02:00

llama-impl.h

Time to fix replace_all (#68 )

2024-09-28 17:59:47 +03:00

llama-sampling.cpp

Merge mainline llama.cpp (#3 )

2024-07-27 07:55:01 +02:00

llama-sampling.h

Merge mainline llama.cpp (#3 )

2024-07-27 07:55:01 +02:00

llama-vocab.cpp

Deepseek V3 support added (#176 )

2025-01-23 18:24:10 +02:00

llama-vocab.h

Merge mainline - Aug 12 2024 (#17 )

2024-08-12 15:14:32 +02:00

llama.cpp

Faster MoE token generation on CUDA (#248 )

2025-03-10 16:16:51 +02:00

unicode-data.cpp

Merge mainline llama.cpp (#3 )

2024-07-27 07:55:01 +02:00

unicode-data.h

Merge mainline llama.cpp (#3 )

2024-07-27 07:55:01 +02:00

unicode.cpp

Deepseek V3 support added (#176 )

2025-01-23 18:24:10 +02:00

unicode.h

Merge mainline llama.cpp (#3 )

2024-07-27 07:55:01 +02:00