< cd ../gallery
[AI & ML]

turboquant

First open-source implementation of Google TurboQuant (ICLR 2026) -- near-optimal KV cache compression for LLM inference. 5x compression with near-zero quality loss.
OnlyTerp
@OnlyTerp
maintainer
★ 76 stars
github
turboquant — preview
turboquant repository preview
# README
First open-source implementation of Google TurboQuant (ICLR 2026) -- near-optimal KV cache compression for LLM inference. 5x compression with near-zero quality loss. Maintained by OnlyTerp on GitHub, where it has earned 76 stars from the community.
It's actively developed around attention, compression, deep-learning, and is a solid reference for anyone building with these tools.
# tags
# install
npm install turboquant
languages
Python100%
last commit2 months ago
licenseMIT
more vLLM repos
$ made-with-vllm
rssllmmadewithwhat

Ask MadeWithWhat

AI answers may contain mistakes — please double-check important details.