MFMA builtins

MFMA builtins#

Matrix Fused Multiply-Add (MFMA) builtins let you issue hardware matrix multiply-accumulate operations directly from HIP device code. This reference covers both dense and sparse MFMA variants on AMD CDNA architectures (AMD Instinct GPUs). Coverage is organized first by operand density – dense or sparse – and then by CDNA generation. The CDNA4 MFMA LDS transpose load builtins are covered on a separate page.