< cd ../gallery
[AI & ML]

LLMKube

Kubernetes operator for self-hosted LLM inference across a heterogeneous GPU fleet: NVIDIA CUDA, AMD Vulkan, and Apple Silicon Metal. Runtimes: llama.cpp, vLLM, TGI, mlx-server. Multi-GPU sharding, model caching, OpenAI-compatible endpoints. Apache-2.0, run across homelab and on-prem fleets, actively developed.
defilantech
@defilantech
maintainer
★ 166 stars
go to website ↗github
llmkube — preview
LLMKube repository preview
# README
Kubernetes operator for self-hosted LLM inference across a heterogeneous GPU fleet: NVIDIA CUDA, AMD Vulkan, and Apple Silicon Metal. Runtimes: llama.cpp, vLLM, TGI, mlx-server. Multi-GPU sharding, model caching, OpenAI-compatible endpoints. Apache-2.0, run across homelab and on-prem fleets, actively developed. Maintained by defilantech on GitHub, where it has earned 166 stars from the community.
It's actively developed around ai, apple-silicon, autoscaling, and is a solid reference for anyone building with these tools.
# tags
# install
npm install LLMKube
languages
Go96.932%
Shell1.722%
Makefile0.467%
HCL0.386%
Python0.155%
last commit1 week ago
licenseApache-2.0
more vLLM repos
$ made-with-vllm
rssllmmadewithwhat

Ask MadeWithWhat

AI answers may contain mistakes — please double-check important details.