FP8-as-Storage GEMM for Ampere
FP8 matmul on an RTX 3090, no H100 required
IMMA-based FP8-as-storage GEMM experiments for Ampere (sm_86 / RTX 3090 Ti). Uses the integer matrix-multiply units to move FP8 data through hardware that was never meant to support it.
Writeup: Backporting FP8 to the RTX 3090