< cd ../gallery
[AI & ML]

tiny-vllm

Build your own high performance LLM inference engine in C++ and CUDA - a smaller version of vLLM
jmaczan
@jmaczan
maintainer
★ 919 stars
github
tiny-vllm — preview
tiny-vllm repository preview
# README
Build your own high performance LLM inference engine in C++ and CUDA - a smaller version of vLLM. Maintained by jmaczan on GitHub, where it has earned 919 stars from the community.
It's actively developed around ai, attention, batching, and is a solid reference for anyone building with these tools.
# tags
# install
npm install tiny-vllm
languages
C++97.758%
Cuda1.808%
C0.192%
CMake0.186%
Shell0.056%
last commit3 weeks ago
licenseApache-2.0
more vLLM repos
$ made-with-vllm
rssllmmadewithwhat

Ask MadeWithWhat

AI answers may contain mistakes — please double-check important details.