Which algorithm is implemented in DGEMM?

Laasner__Raul — Tue, 26 Sep 2017 18:52:34 GMT

Interestingly, I've been unable to find an answer to this simple question. What is the algorithm that is used for matrix-matrix multiplications (e.g., DGEMM) in MKL? Is is classical (O(N^3)), Strassen (O(N^2.7)), or something else? Thanks.

Hi Raul,

Zhen_Z_Intel — Wed, 27 Sep 2017 02:07:00 GMT

Hi Raul,

I am afraid BLAS standard gemm uses classical O(N³), for algorithm design, you could follow Netlib gemm source code. Intel MKL optimized BLAS routines with SIMD instruction sets, do some work to fit data into the caches enabling contiguous, aligned accesses.

Here's another algorithm for matrix matrix multiplication, call 3M. It split a complex matrix into two matrices, performs 3 GEMM and 4 matrix additions. For other algorithm, like Winograd which implemented for NN convolution kernel in MKL-DNN.

topic Which algorithm is implemented in DGEMM? in Intel® oneAPI Math Kernel Library

Which algorithm is implemented in DGEMM?

Hi Raul,