< cd ../gallery
[AI & ML]
blackwell-geforce-nvfp4-gemm
NVFP4 inference on Blackwell GeForce (RTX 5090/5080/5070 Ti/RTX PRO 6000) — SM120 patches for vLLM + FlashInfer + CUTLASS. 175 tok/s on Qwen3.6-35B MoE.
@lna-lab
maintainer
★ 20 stars
# README
NVFP4 inference on Blackwell GeForce (RTX 5090/5080/5070 Ti/RTX PRO 6000) — SM120 patches for vLLM + FlashInfer + CUTLASS. 175 tok/s on Qwen3.6-35B MoE. Maintained by lna-lab on GitHub, where it has earned 20 stars from the community.
It's actively developed around blackwell, cutlass, flashinfer, and is a solid reference for anyone building with these tools.
# tags
# install
npm install blackwell-geforce-nvfp4-gemmmore vLLM repos