carlushuang
|
3d15f364b3
|
[CK_TILE] optimize moe-sorting kernel (#1771)
* opt moe sorting
* remove commented code
|
2024-12-23 10:59:02 +08:00 |
|
Xu, Shengnan
|
f57d720c67
|
added moe interleaving pipeline (#1712)
* added moe interleaving pipeline
* remove redundant code
* formater
---------
Co-authored-by: root <root@hjbog-srdc-14.amd.com>
|
2024-12-15 20:13:10 +08:00 |
|
carlushuang
|
440e28b08f
|
[CK_TILE] fused-moe first version (#1634)
* moe pipeline
* update code
* compile OK
* update
* update cpu reference
* update pipeline_gemm0
* compiler ok
* update pipeline
* rename to ex pipeline
* block-asm
* update
* update
* update first gemm ok
* compute correct
* update file structure
* update README
* update
* update
* update code
* update API
* return unsupport case
* add comment
* update readme
* update
* uncomment
* update
* fix build err
---------
Co-authored-by: valarLip <340077269@qq.com>
|
2024-11-26 11:14:56 +08:00 |
|
carlushuang
|
36c7ce4e0e
|
[CK_TILE]Moe update index (#1672)
* update MOCK_ID for moe-sorting
* add moe-smoothquant
* update a comment
* fix format
* hot fix
* update topk in overflow case
* update comments
* update bf16 cvt
---------
Co-authored-by: valarLip <340077269@qq.com>
|
2024-11-25 13:12:35 +08:00 |
|
dummycoderfe
|
bec6fbc65f
|
Ck tile/moe sorting (#1624)
* add moe_sorting & check ok
* fix comments & typo
* Run remod.py under include/ck_tile & example/ck_tile directories
* format codes
* fix output ci check bug
* fix moe sorting readme and error commit file
* use magiv div to accelerate compute
* add an loop unroll for moe lds ops
* add extblocksnel to set zeros for moebufs
* [Ck_tile] moe set zero run ok, add size check and fix ref check
* [Ck_tile]fix moe_sorting fuse set_zero remod
* [Ck_tile] change name style, fix zero buffer size err, change folder
* [Ck_tile] moe_sorting: fix name style
* [Ck_tile] moe_sorting, remove useless params in traits
* [Ck_tile] change outputtile cnt * unit_size; change output buf alloc
---------
Co-authored-by: dummycoderfe <noplydummmycoder@163.com>
Co-authored-by: Po Yen, Chen <PoYen.Chen@amd.com>
Co-authored-by: carlushuang <carlus.huang@amd.com>
|
2024-11-09 17:57:27 +08:00 |
|