mscclpp

mirror of https://github.com/microsoft/mscclpp.git synced 2026-07-12 18:27:10 +00:00

Files

Binyang Li 8896cd909a Add ROCm FP8 E4M3B15 support (#774 )

## Summary

Add ROCm (gfx942) support for the FP8 E4M3B15 data type, including
optimized conversion routines between FP8 E4M3B15 and FP16/FP32 using
inline assembly.

Extends the allpair packet and fullmesh allreduce kernels to support
higher-precision accumulation (e.g., FP16/FP32) when reducing FP8 data,
improving numerical accuracy.

Adds Python tests to verify that higher-precision accumulation is at
least as accurate as native FP8 accumulation across all algorithm
variants.

---------

Co-authored-by: Claude Opus 4.6 <noreply@anthropic.com>
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>

2026-04-08 09:53:45 -07:00

ext

Refactor algo selection logic and introduce symmetric_memory env (#741 )

2026-02-12 19:06:18 -08:00

algorithm.cpp

Support E4M3B15 datatype (#765 )

2026-04-07 13:37:02 -07:00

CMakeLists.txt

Add ROCm FP8 E4M3B15 support (#774 )

2026-04-08 09:53:45 -07:00

core_py.cpp

Support E4M3B15 datatype (#765 )