< cd ../gallery
[AI & ML]
multi-turboquant
Unified KV-cache compression for LLM inference: 12 Python-native methods, Debian-tested isolated add-ons, Godzilla KVarN/TriAttention, exact Godzilla/Gigatoken profiles, CUDA weight sharing, and multi-GPU planning.
@aivrar
maintainer
★ 25 stars
# README
Unified KV-cache compression for LLM inference: 12 Python-native methods, Debian-tested isolated add-ons, Godzilla KVarN/TriAttention, exact Godzilla/Gigatoken profiles, CUDA weight sharing, and multi-GPU planning. Maintained by aivrar on GitHub, where it has earned 25 stars from the community.
It's actively developed around attention, compression, cuda, and is a solid reference for anyone building with these tools.
# tags
# install
npm install multi-turboquantmore vLLM repos