Chao Liu
|
db876ea7ec
|
adding implicit gemm v4 (nchw, kcyx)
[ROCm/composable_kernel commit: b2439ec9dd]
|
2019-05-30 17:50:49 -05:00 |
|
Chao Liu
|
979dc4da2e
|
adding implicit gemm v3
[ROCm/composable_kernel commit: 8a4b59785b]
|
2019-05-22 19:39:56 -05:00 |
|
Chao Liu
|
45e1ad4dea
|
adding ConstantMergedTensorDescriptor, refactering ConstantTensorDescriptor, Sequence
[ROCm/composable_kernel commit: acd7082fe1]
|
2019-05-21 16:17:58 -05:00 |
|
Chao Liu
|
ffd172378a
|
adding implicit gemm v3
[ROCm/composable_kernel commit: 5e5c27a63b]
|
2019-05-16 13:22:40 -05:00 |
|
Chao Liu
|
ac7741cc7c
|
adding implicit gemm v3
[ROCm/composable_kernel commit: b7d052459d]
|
2019-05-15 09:58:17 -05:00 |
|
Chao Liu
|
04e99df5df
|
tuning on vega 20
[ROCm/composable_kernel commit: 2603bb0fe3]
|
2019-04-25 17:28:59 -05:00 |
|
Chao Liu
|
868068eee1
|
implicit gemm v1r3 nchw_cyxk_nkhw
[ROCm/composable_kernel commit: a903146427]
|
2019-04-25 15:14:39 -05:00 |
|
Chao Liu
|
21988c32b4
|
added implicit gemm v1r3 lds_double_buffer NCHW * CYXK = KNHW, reworked static functionals
[ROCm/composable_kernel commit: 569ad66e2a]
|
2019-04-23 17:51:14 -05:00 |
|
Chao Liu
|
f367814ef7
|
added GridwiseConvolutionImplicitGemm_v1r2_nchw_cyxk_khwn
[ROCm/composable_kernel commit: 5ce19234a4]
|
2019-04-19 14:22:02 -05:00 |
|
Chao Liu
|
d0244d3a51
|
implicit gemm v1r2: adding support for nchw
[ROCm/composable_kernel commit: 19f17df47a]
|
2019-04-18 11:49:09 -05:00 |
|
Chao Liu
|
482e5e9293
|
refactor ConstantTensorDescriptor and functional
[ROCm/composable_kernel commit: 17f3d2d4bc]
|
2019-04-16 17:36:18 -05:00 |
|
Chao Liu
|
8b7eafe959
|
implicit gemm v1r2: only load 1d filter
[ROCm/composable_kernel commit: 00899f191b]
|
2019-04-13 11:19:17 -05:00 |
|
Chao Liu
|
2fc34a9169
|
tuned implicit gemm v1 for 3x3 on AMD to 82%. Fixed a bug in 4d tensor blockwise copy.
[ROCm/composable_kernel commit: 96ee9571e2]
|
2019-04-10 18:10:18 -05:00 |
|
Chao Liu
|
75ca00f748
|
tidy yp
[ROCm/composable_kernel commit: 471830a052]
|
2019-04-09 18:07:36 -05:00 |
|
Chao Liu
|
709d9581bf
|
added implicit_gemm_v1 lds double_buffer
[ROCm/composable_kernel commit: cc0fa73acd]
|
2019-04-08 14:11:55 -05:00 |
|
Chao Liu
|
8cb0a5b05f
|
tidy up
[ROCm/composable_kernel commit: 268d1c717c]
|
2019-04-08 10:48:29 -05:00 |
|
Chao Liu
|
d430879858
|
debugging implicit gemm v1: use 10d tensor output
[ROCm/composable_kernel commit: c9fa46af0b]
|
2019-04-08 10:27:32 -05:00 |
|
Chao Liu
|
cd883e7581
|
experimenting
[ROCm/composable_kernel commit: 766b0a9eaf]
|
2019-03-24 12:09:57 -05:00 |
|
Chao Liu
|
1f925812b2
|
hip build
[ROCm/composable_kernel commit: 8c923db423]
|
2019-03-22 14:22:58 -05:00 |
|
Chao Liu
|
732984e63b
|
adding fp16 direct that reads pre-vectorized data
[ROCm/composable_kernel commit: 79d9b1084b]
|
2019-03-18 18:16:02 -05:00 |
|
Chao Liu
|
5ba0f64087
|
adding fp16 direct that reads pre-vectorized data
[ROCm/composable_kernel commit: 4f0fc72e91]
|
2019-03-18 15:03:17 -05:00 |
|
Chao Liu
|
da50d65ba0
|
refactoring block copy
[ROCm/composable_kernel commit: 03eef73c5b]
|
2019-03-17 15:36:38 -05:00 |
|
Chao Liu
|
e3ab560c50
|
refactor
[ROCm/composable_kernel commit: 04c5527d07]
|
2019-03-04 17:09:20 -06:00 |
|
Chao Liu
|
13388cef47
|
add anther verision of batch gemm
[ROCm/composable_kernel commit: 1cb9885058]
|
2019-02-17 01:50:57 -06:00 |
|
Chao Liu
|
c0baa18a3f
|
change file extension to hip.hpp and hip.cpp
[ROCm/composable_kernel commit: b2888adfbe]
|
2019-02-15 02:13:21 -06:00 |
|