< cd ../gallery
[AI & ML]
tangram
A high-throughput LLM serving engine with non-uniform KV cache compression, built on vLLM
@aiha-lab
maintainer
★ 13 stars
# README
A high-throughput LLM serving engine with non-uniform KV cache compression, built on vLLM. Maintained by aiha-lab on GitHub, where it has earned 13 stars from the community.
It's actively developed around inference-optimization, kv-cache-compression, llm-inference, and is a solid reference for anyone building with these tools.
# tags
# install
npm install tangrammore vLLM repos