< cd ../gallery
[AI & ML]

auto-round

A SOTA quantization algorithm for high-accuracy low-bit LLM inference, seamlessly optimized for CPU/XPU/CUDA, with multi-datatype support and full compatibility with vLLM, SGLang, and Transformers.
intel
@intel
maintainer
★ 1.5k stars
github
auto-round — preview
auto-round repository preview
# README
A SOTA quantization algorithm for high-accuracy low-bit LLM inference, seamlessly optimized for CPU/XPU/CUDA, with multi-datatype support and full compatibility with vLLM, SGLang, and Transformers. Maintained by intel on GitHub, where it has earned 1,519 stars from the community.
It's actively developed around diffusers, gguf, int4, and is a solid reference for anyone building with these tools.
# tags
# install
npm install auto-round
languages
Python69.324%
C++30.015%
Shell0.424%
CMake0.207%
C0.029%
last commit1 week ago
licenseApache-2.0
more vLLM repos
$ made-with-vllm
rssllmmadewithwhat

Ask MadeWithWhat

AI answers may contain mistakes — please double-check important details.