85.3 GFlops: Optimizing FP32 Matrix Multiplication on a Single AMD Zen 3 Core
Summary
This article presents a systematic study of optimizing FP32 matrix multiplication (GEMM) for a single AMD Zen 3 core using AVX2/FMA in C++ intrinsics. It explores 28 configurations, identifies MX24 as the best performer with 85.30 GFLOPS (63.5% of theoretical peak), and details the impact of techniques like cache tiling, register blocking, FMA chaining, packing strategies, alignment, prefetching, and stores. The work also documents methodology, validation against a reference, and practical guidance for reproducing the results.