Details: - Added SIMD code - Processing 5 rows at a time in SIMD loop to improve performance AMD-Internal: [CPUPL-1054] Change-Id: I2ac93f25895dccfc42e14be0689e6d4e655d6a0a
b426f9e