< cd ../gallery
[AI & ML]

multi-turboquant

Unified KV-cache compression for LLM inference: 12 Python-native methods, Debian-tested isolated add-ons, Godzilla KVarN/TriAttention, exact Godzilla/Gigatoken profiles, CUDA weight sharing, and multi-GPU planning.
aivrar
@aivrar
maintainer
★ 25 stars
github
multi-turboquant — preview
multi-turboquant repository preview
# README
Unified KV-cache compression for LLM inference: 12 Python-native methods, Debian-tested isolated add-ons, Godzilla KVarN/TriAttention, exact Godzilla/Gigatoken profiles, CUDA weight sharing, and multi-GPU planning. Maintained by aivrar on GitHub, where it has earned 25 stars from the community.
It's actively developed around attention, compression, cuda, and is a solid reference for anyone building with these tools.
# tags
# install
npm install multi-turboquant
languages
Python100%
last commit4w ago
licenseMIT
more vLLM repos

Ask MadeWithWhat

AI answers may contain mistakes — please double-check important details.