面向 8×RTX 4090(SM89)的 DeepSeek-V4-Flash-0731 MXFP4 SGLang 实验分支
-
Updated
Aug 11, 2026 - Python
面向 8×RTX 4090(SM89)的 DeepSeek-V4-Flash-0731 MXFP4 SGLang 实验分支
Feeding the Tensor Cores: a dense FP16 GEMM for NVIDIA Ada (sm_89) at 96.5% of cuBLAS, and what it teaches about how each GPU generation handles async copy and Tensor Core issue.
Add a description, image, and links to the sm89 topic page so that developers can more easily learn about it.
To associate your repository with the sm89 topic, visit your repo's landing page and select "manage topics."