< cd ../gallery
[AI & ML]
time-to-first-token
A 10-week, 30-minutes-a-day roadmap for LLM inference serving and optimization. vLLM, SGLang, quantization, speculative decoding, benchmarking.
@patchy631
maintainer
★ 296 stars
# README
A 10-week, 30-minutes-a-day roadmap for LLM inference serving and optimization. vLLM, SGLang, quantization, speculative decoding, benchmarking. Maintained by patchy631 on GitHub, where it has earned 296 stars from the community.
It's actively developed around learning-resources, llm, llm-inference, and is a solid reference for anyone building with these tools.
# tags
# install
npm install time-to-first-tokenmore vLLM repos