Files
ik_llama.cpp/common
Kawrakow b66cecca45 Fused FFN_UP+FFN_GATE op (#741)
* Fused up+gate+unary for regular (not MoE) FFN - CPU

* WIP CUDA

* Seems to be working on CUDA

For a dense model we get 2-3% speedup for PP and ~0.6% for TG.

* Add command line option

This time the option is ON by default, and one needs to turn it
off via -no-fug or --no-fused-up-gate

---------

Co-authored-by: Iwan Kawrakow <iwan.kawrakow@gmail.com>
2025-08-31 18:16:36 +03:00
..
2024-07-27 07:55:01 +02:00
2025-08-31 18:16:36 +03:00
2025-08-31 18:16:36 +03:00
2024-07-27 07:55:01 +02:00
2025-08-09 12:50:30 +00:00
2024-07-27 07:55:01 +02:00
2023-11-13 14:16:23 +02:00