< cd ../gallery
[AI & ML]
vllm-windows-build
Native Windows vLLM 0.26.0: CPython 3.13, CUDA 12.8, SM 7.5/8.6/8.9/12.0 for RTX 20/30/40/50, OpenAI-compatible serving, Triton/FlashAttention, 10 KV-cache formats, Multi-TurboQuant, and experimental CPU/RAM/NVMe prompt-KV offload. No WSL or Docker.
@aivrar
maintainer
★ 36 stars
# README
Native Windows vLLM 0.26.0: CPython 3.13, CUDA 12.8, SM 7.5/8.6/8.9/12.0 for RTX 20/30/40/50, OpenAI-compatible serving, Triton/FlashAttention, 10 KV-cache formats, Multi-TurboQuant, and experimental CPU/RAM/NVMe prompt-KV offload. No WSL or Docker. Maintained by aivrar on GitHub, where it has earned 36 stars from the community.
It's actively developed around cuda, gpu, llm, and is a solid reference for anyone building with these tools.
# tags
# install
npm install vllm-windows-buildmore vLLM repos