Files
composable_kernel/include/ck/tensor_operation/gpu/device
Mingtao Gu 7998ae8969 [CK] Mxfp4 moe blockscale buf2lds version support (#2455)
* change cshuffle size

* added mxfp4 moe async buffer loading without B preshuffle

* added mx moe B shuffling + scale shuffling (async loads)

* minor fix

---------

Co-authored-by: mtgu0705 <mtgu@amd.com>
2025-07-06 15:42:00 +08:00
..
2024-03-08 17:11:51 -08:00
2025-03-10 11:16:44 +08:00
2023-08-15 02:25:28 +08:00
2024-06-25 16:37:35 -05:00