ik_llama.cpp/ggml-cuda/common.cuh at ea3b0590ee33d3573eb8ef76f88cc60f36d2a38d

mirror of https://github.com/ikawrakow/ik_llama.cpp.git synced 2026-04-28 18:32:04 +00:00

Files

Johannes Gäßler dc685be466 CUDA: add FP32 FlashAttention vector kernel (#7188 )

* CUDA: add FP32 FlashAttention vector kernel

* fixup! CUDA: add FP32 FlashAttention vector kernel

* fixup! fixup! CUDA: add FP32 FlashAttention vector kernel

* fixup! fixup! fixup! CUDA: add FP32 FlashAttention vector kernel

2024-05-12 19:40:45 +02:00

23 KiB

Raw Blame History

View Raw

23 KiB Raw Blame History

23 KiB

Raw Blame History