< cd ../gallery
[AI & ML]
auto-round
A SOTA quantization algorithm for high-accuracy low-bit LLM inference, seamlessly optimized for CPU/XPU/CUDA, with multi-datatype support and full compatibility with vLLM, SGLang, and Transformers.
@intel
maintainer
★ 1.5k stars
# README
A SOTA quantization algorithm for high-accuracy low-bit LLM inference, seamlessly optimized for CPU/XPU/CUDA, with multi-datatype support and full compatibility with vLLM, SGLang, and Transformers. Maintained by intel on GitHub, where it has earned 1,519 stars from the community.
It's actively developed around diffusers, gguf, int4, and is a solid reference for anyone building with these tools.
# tags
# install
npm install auto-roundmore vLLM repos