* fix(grouped_gemm): numerical errors on gfx950 by correctly calculating the tail num * WIP: add temp config to stress test numerical error correction * refactor: remove comments [ROCm/composable_kernel commit: db79fad16f]
db79fad16f