< cd ../gallery
[AI & ML]

Lvllm

LvLLM is a special NUMA extension of vllm that makes full use of CPU and memory resources, reduces GPU memory requirements, and features an efficient GPU parallel and NUMA parallel architecture, supporting hybrid inference for MOE large models.
guqiong96
@guqiong96
maintainer
★ 384 stars
github
lvllm — preview
Lvllm repository preview
# README
LvLLM is a special NUMA extension of vllm that makes full use of CPU and memory resources, reduces GPU memory requirements, and features an efficient GPU parallel and NUMA parallel architecture, supporting hybrid inference for MOE large models. Maintained by guqiong96 on GitHub, where it has earned 384 stars from the community.
It's actively developed around cpu, decode, gpu, and is a solid reference for anyone building with these tools.
# tags
# install
npm install Lvllm
languages
Python84.456%
Rust5.523%
Cuda4.957%
C++3.596%
Shell1.001%
CMake0.279%
HCL0.046%
C0.022%
Jinja0.014%
last commit1 week ago
licenseApache-2.0
more vLLM repos
$ made-with-vllm
rssllmmadewithwhat

Ask MadeWithWhat

AI answers may contain mistakes — please double-check important details.