< cd ../gallery
[AI & ML]

multi-turboquant

Unified KV cache compression for LLM inference — TurboQuant, IsoQuant, PlanarQuant, TriAttention. 10 methods, GPU-validated, multi-GPU planner. Compress KV cache 5-80x to run bigger models, longer context, more agents on your GPU.
aivrar
@aivrar
maintainer
★ 24 stars
github
multi-turboquant — preview
multi-turboquant repository preview
# README
Unified KV cache compression for LLM inference — TurboQuant, IsoQuant, PlanarQuant, TriAttention. 10 methods, GPU-validated, multi-GPU planner. Compress KV cache 5-80x to run bigger models, longer context, more agents on your GPU. Maintained by aivrar on GitHub, where it has earned 24 stars from the community.
It's actively developed around attention, compression, cuda, and is a solid reference for anyone building with these tools.
# tags
# install
npm install multi-turboquant
languages
Python99.118%
Shell0.882%
last commit2 weeks ago
licenseMIT
more vLLM repos
$ made-with-vllm
rssllmmadewithwhat

Ask MadeWithWhat

AI answers may contain mistakes — please double-check important details.