< cd ../gallery
[AI & ML]

nexusquant

Training-free KV cache compression via E8 lattice VQ. 2-bit KV that preserves retrieval (30/30 NIAH vs TurboQuant 0/30). Calibration-free, 9 architectures validated.
jagmarques
@jagmarques
maintainer
★ 25 stars
github
nexusquant — preview
nexusquant repository preview
# README
Training-free KV cache compression via E8 lattice VQ. 2-bit KV that preserves retrieval (30/30 NIAH vs TurboQuant 0/30). Calibration-free, 9 architectures validated. Maintained by jagmarques on GitHub, where it has earned 25 stars from the community.
It's actively developed around attention, compression, e8-lattice, and is a solid reference for anyone building with these tools.
# tags
# install
npm install nexusquant
languages
Python100%
last commit1 week ago
license
more vLLM repos
$ made-with-vllm
rssllmmadewithwhat

Ask MadeWithWhat

AI answers may contain mistakes — please double-check important details.