BF16 Component-Product Emulation of FP32 and FP64 GEMM on Intel AMX
Bing Cui, Yu Liu
Abstract
Modern CPUs increasingly integrate high-throughput matrix engines optimized for low-precision AI workloads, while many scientific computing applications still rely on FP32 and FP64 GEMM to meet their numerical accuracy requirements. This mismatch motivates an algorithmic bridge that uses low-precision matrix products to emulate higher-precision GEMM. This paper presents a CPU-oriented method based on Intel Advanced Matrix Extensions (AMX) and BF16 matrix products. For FP32, each operand is decomposed into three BF16 components and six selected component products are evaluated, targeting FP32-level accuracy relative to oneMKL SGEMM without claiming elementwise or bitwise identity. For FP64 inputs within the supported BF16 exponent range, the method uses a simplified fixed six-slice Ozaki decomposition. Each retained BF16 product is first produced in FP32, then widened and accumulated in FP64. Four product-count settings retain 6, 10, 15, or 21 component products, exposing the accuracy--performance tradeoff relative to oneMKL DGEMM. The implementation combines precomputed packed component buffers, VNNI-packed B panels, and an FP32 tile-resident operand-reuse schedule. On the tested square matrices, AMX-FP32 exceeds oneMKL SGEMM throughput. For AMX-FP64, low-product-count variants can exceed DGEMM at sufficiently large orders, while retaining more products improves accuracy at additional cost.
Create a lesson
Related papers
The Art of Closed-Formula Defaults: Search-Free Code Generation for Tensor Operators
Paolo D'Alberto, Ashish Sirasao
Geometric Function Atlas: certified computing for geometric function theory in Python
Kishan Gurumurthy, Pushparaj Devadiga, Prasanna Devadiga et al.
Gradient Reconstruction in Lattice Boltzmann Methods for Systems of Conservation Laws
Adrian Kummerländer, Fedor Bukreev, Mathias J. Krause
GRADSOLVE: fast exact gradients for ODE ensembles on GPUs
Alessio Spurio Mancini
Performance Evaluation of Fast Fourier Transforms on Emerging RISC-V Hardware with Vector Extension Support
Daniel Seibel, Kaveh Haghighi Mood, Jayesh Badwaik et al.
The Pauli Lightcone: Information-Theoretic Error Mitigation Beyond the Autocorrelation
Paolo D'Alberto