< cd ../gallery
[AI & ML]
multi-turboquant
Unified KV cache compression for LLM inference — TurboQuant, IsoQuant, PlanarQuant, TriAttention. 10 methods, GPU-validated, multi-GPU planner. Compress KV cache 5-80x to run bigger models, longer context, more agents on your GPU.
@aivrar
maintainer
★ 24 stars
# README
Unified KV cache compression for LLM inference — TurboQuant, IsoQuant, PlanarQuant, TriAttention. 10 methods, GPU-validated, multi-GPU planner. Compress KV cache 5-80x to run bigger models, longer context, more agents on your GPU. Maintained by aivrar on GitHub, where it has earned 24 stars from the community.
It's actively developed around attention, compression, cuda, and is a solid reference for anyone building with these tools.
# tags
# install
npm install multi-turboquantmore vLLM repos