< cd ../gallery
[AI & ML]

blackwell-geforce-nvfp4-gemm

NVFP4 inference on Blackwell GeForce (RTX 5090/5080/5070 Ti/RTX PRO 6000) — SM120 patches for vLLM + FlashInfer + CUTLASS. 175 tok/s on Qwen3.6-35B MoE.
lna-lab
@lna-lab
maintainer
★ 20 stars
github
blackwell-geforce-nvfp4-gemm — preview
blackwell-geforce-nvfp4-gemm repository preview
# README
NVFP4 inference on Blackwell GeForce (RTX 5090/5080/5070 Ti/RTX PRO 6000) — SM120 patches for vLLM + FlashInfer + CUTLASS. 175 tok/s on Qwen3.6-35B MoE. Maintained by lna-lab on GitHub, where it has earned 20 stars from the community.
It's actively developed around blackwell, cutlass, flashinfer, and is a solid reference for anyone building with these tools.
# tags
# install
npm install blackwell-geforce-nvfp4-gemm
languages
Python100%
last commit3 months ago
license
more vLLM repos
$ made-with-vllm
rssllmmadewithwhat

Ask MadeWithWhat

AI answers may contain mistakes — please double-check important details.