< cd ../gallery
[AI & ML]
nexusquant
Training-free KV cache compression via E8 lattice VQ. 2-bit KV that preserves retrieval (30/30 NIAH vs TurboQuant 0/30). Calibration-free, 9 architectures validated.
@jagmarques
maintainer
★ 25 stars
# README
Training-free KV cache compression via E8 lattice VQ. 2-bit KV that preserves retrieval (30/30 NIAH vs TurboQuant 0/30). Calibration-free, 9 architectures validated. Maintained by jagmarques on GitHub, where it has earned 25 stars from the community.
It's actively developed around attention, compression, e8-lattice, and is a solid reference for anyone building with these tools.
# tags
# install
npm install nexusquantmore vLLM repos