< cd ../gallery
[AI & ML]
tiny-vllm
Build your own high performance LLM inference engine in C++ and CUDA - a smaller version of vLLM
@jmaczan
maintainer
★ 994 stars
# README
Build your own high performance LLM inference engine in C++ and CUDA - a smaller version of vLLM. Maintained by jmaczan on GitHub, where it has earned 994 stars from the community.
It's actively developed around ai, hpc, llm, and is a solid reference for anyone building with these tools.
# tags
# install
npm install tiny-vllmmore vLLM repos