Blackwell-optimized llama.cpp bundle with DFlash2, MXFP6/MXFP8/NVFP4, TurboQuant KV, and GPU-resident speculative handoff.
-
Updated
Aug 23, 2026 - C++
Blackwell-optimized llama.cpp bundle with DFlash2, MXFP6/MXFP8/NVFP4, TurboQuant KV, and GPU-resident speculative handoff.
Add a description, image, and links to the mxfp6 topic page so that developers can more easily learn about it.
To associate your repository with the mxfp6 topic, visit your repo's landing page and select "manage topics."