< cd ../gallery
[AI & ML]

Efficient-LLM-Inference-Serving-Systems

Why is LLM inference slow — and how do you make it fast? A hands-on, first-principles course: roofline → KV cache → quantization → parallelism → vLLM/SGLang, with GPU labs on open models.
jiahongsigma
@jiahongsigma
maintainer
★ 19 stars
github
efficient-llm-inference-serving-systems — preview
Efficient-LLM-Inference-Serving-Systems repository preview
# README
Why is LLM inference slow — and how do you make it fast? A hands-on, first-principles course: roofline → KV cache → quantization → parallelism → vLLM/SGLang, with GPU labs on open models. Maintained by jiahongsigma on GitHub, where it has earned 19 stars from the community.
It's actively developed around cuda, deep-learning, flash-attention, and is a solid reference for anyone building with these tools.
# tags
# install
npm install Efficient-LLM-Inference-Serving-Systems
languages
Python98.526%
Shell1.474%
last commit3 weeks ago
licenseMIT
more vLLM repos
$ made-with-vllm
rssllmmadewithwhat

Ask MadeWithWhat

AI answers may contain mistakes — please double-check important details.