DigiNews

Tech Watch by Johan Denoyer

← Back to articles

85.3 GFlops: Optimizing FP32 Matrix Multiplication on a Single AMD Zen 3 Core

Quality: 8/10 Relevance: 9/10

Summary

This article presents a systematic study of optimizing FP32 matrix multiplication (GEMM) for a single AMD Zen 3 core using AVX2/FMA in C++ intrinsics. It explores 28 configurations, identifies MX24 as the best performer with 85.30 GFLOPS (63.5% of theoretical peak), and details the impact of techniques like cache tiling, register blocking, FMA chaining, packing strategies, alignment, prefetching, and stores. The work also documents methodology, validation against a reference, and practical guidance for reproducing the results.

🚀 Service construit par Johan Denoyer