< cd ../gallery
[AI & ML]

tangram

A high-throughput LLM serving engine with non-uniform KV cache compression, built on vLLM
aiha-lab
@aiha-lab
maintainer
★ 13 stars
github
tangram — preview
tangram repository preview
# README
A high-throughput LLM serving engine with non-uniform KV cache compression, built on vLLM. Maintained by aiha-lab on GitHub, where it has earned 13 stars from the community.
It's actively developed around inference-optimization, kv-cache-compression, llm-inference, and is a solid reference for anyone building with these tools.
# tags
# install
npm install tangram
languages
C++4%
last commit4w ago
licenseApache-2.0
more vLLM repos

Ask MadeWithWhat

AI answers may contain mistakes — please double-check important details.