# Made with vLLM — full catalog > A curated, daily-updated gallery of the best open-source projects built with vLLM, ranked by GitHub stars. Discover dashboards, UI kits, e-commerce, blogs and dev tools. ## About - Gallery: https://madewithwhat.net/vllm/ - Curated summary: https://madewithwhat.net/vllm/llms.txt - Projects indexed: 289 - Data source: GitHub (refreshed daily) - Last scraped: 2026-07-22T09:29:08.088214+00:00 ## AI & ML (244) - [FunASR](https://madewithwhat.net/vllm/project/funasr/): Industrial-grade speech recognition toolkit: 170x realtime, 50+ languages, speaker diarization, emotion detection, streaming, and OpenAI-compatible API. (19,219 stars, MIT) - [llama-cookbook](https://madewithwhat.net/vllm/project/llama-cookbook/): Welcome to the Llama Cookbook! This is your go to guide for Building with Llama: Getting started with Inference, Fine-Tuning, RAG. We also show you how to solve end to end problems using Llama model family and using them on various provider services (18,401 stars, MIT) - [AI-Research-SKILLs](https://madewithwhat.net/vllm/project/ai-research-skills/): Comprehensive open-source library of AI research and engineering skills for any AI model. Package the skills and your claude code/codex/gemini agent will be an AI research agent with full horsepower. Maintained by Orchestra Research. (10,689 stars, MIT) - [LMCache](https://madewithwhat.net/vllm/project/lmcache/): LMCache: Supercharge Your LLM with the Fastest KV Cache Layer (10,541 stars, Apache-2.0) - [OpenRLHF](https://madewithwhat.net/vllm/project/openrlhf/): An Easy-to-use, Scalable and High-performance Agentic RL Framework based on Ray (PPO & DAPO & REINFORCE++ & VLM & TIS & vLLM & Ray & Async RL) (9,785 stars, Apache-2.0) - [inference](https://madewithwhat.net/vllm/project/inference/): Swap GPT for any LLM by changing a single line of code. Xinference lets you run open-source, speech, and multimodal models on cloud, on-prem, or your laptop — all through one unified, production-ready inference API. (9,429 stars, Apache-2.0) - [dynamo](https://madewithwhat.net/vllm/project/dynamo/): A Datacenter Scale Distributed Inference Serving Framework (7,479 stars) - [Mooncake](https://madewithwhat.net/vllm/project/mooncake/): Mooncake is the serving platform for Kimi, a leading LLM service provided by Moonshot AI. (5,816 stars, Apache-2.0) - [kserve](https://madewithwhat.net/vllm/project/kserve/): Standardized Distributed Generative and Predictive AI Inference Platform for Scalable, Multi-Framework Deployment on Kubernetes (5,681 stars, Apache-2.0) - [UltraRAG](https://madewithwhat.net/vllm/project/ultrarag/): A Low-Code MCP Framework for Building Complex and Innovative RAG Pipelines (5,643 stars, Apache-2.0) - [gpustack](https://madewithwhat.net/vllm/project/gpustack/): A GPU cluster manager for high-performance AI model serving (vLLM, SGLang) and on-demand SSH-accessible GPU instances. (5,317 stars, Apache-2.0) - [sparrow](https://madewithwhat.net/vllm/project/sparrow/): Structured data extraction, instruction calling and agentic workflows with ML, LLM and Vision LLM (5,181 stars, GPL-3.0) - [llama-swap](https://madewithwhat.net/vllm/project/llama-swap/): Reliable model swapping for any local OpenAI/Anthropic compatible server - llama.cpp, vllm, etc (4,992 stars, MIT) - [semantic-router](https://madewithwhat.net/vllm/project/semantic-router/): Intelligent Mixture-of-Models Router for Efficient Heterogeneous LLMs Inference (4,962 stars, Apache-2.0) - [tiny-llm](https://madewithwhat.net/vllm/project/tiny-llm/): A course of learning LLM inference serving on Apple Silicon for systems engineers: build a tiny vLLM + Qwen. (4,361 stars, Apache-2.0) - [LazyLLM](https://madewithwhat.net/vllm/project/lazyllm/): Easiest and laziest way for building multi-agent LLMs applications. (3,853 stars, Apache-2.0) - [FastDeploy](https://madewithwhat.net/vllm/project/fastdeploy/): High-performance Inference and Deployment Toolkit for LLMs and VLMs based on PaddlePaddle (3,702 stars, Apache-2.0) - [cascadeflow](https://madewithwhat.net/vllm/project/cascadeflow/): Cascading runtime for AI agents. Optimize cost, latency, quality, and policy decisions inside the agent loop. (3,295 stars, MIT) - [Rapid-MLX](https://madewithwhat.net/vllm/project/rapid-mlx/): The fastest local AI engine for Apple Silicon. 4.2x faster than Ollama, 0.08s cached TTFT, 100% tool calling. 17 tool parsers, prompt cache, reasoning separation, cloud routing. Drop-in OpenAI replacement. Works with Claude Code, Cursor, Aider. (3,262 stars, Apache-2.0) - [ramalama](https://madewithwhat.net/vllm/project/ramalama/): RamaLama is an open-source developer tool that simplifies the local serving of AI models from any source and facilitates their use for inference in production, all through the familiar language of containers. (2,956 stars, MIT) - [vllm-ascend](https://madewithwhat.net/vllm/project/vllm-ascend/): Community maintained hardware plugin for vLLM on Ascend (2,496 stars, Apache-2.0) - [auto-round](https://madewithwhat.net/vllm/project/auto-round/): A SOTA quantization algorithm for high-accuracy low-bit LLM inference, seamlessly optimized for CPU/XPU/CUDA, with multi-datatype support and full compatibility with vLLM, SGLang, and Transformers. (1,519 stars, Apache-2.0) - [vllm-mlx](https://madewithwhat.net/vllm/project/vllm-mlx/): OpenAI and Anthropic compatible server for Apple Silicon. Run LLMs and vision-language models (Llama, Qwen-VL, LLaVA) with continuous batching, MCP tool calling, and multimodal support. Native MLX backend, 400+ tok/s. Works with Claude Code. (1,428 stars, Apache-2.0) - [local-studio](https://madewithwhat.net/vllm/project/local-studio/): Control panel for VLLM, Sglang, llama.cpp, exllamav3 (1,359 stars, Apache-2.0) - [InferenceX](https://madewithwhat.net/vllm/project/inferencex/): Open Source Continuous Inference Benchmark Research Platform — Kimi K2.7-Code, MiniMax M3, DeepSeekv4, GLM5 - GB200 NVL72 vs MI355X vs B200 vs GB300 NVL72 & soon™ TPUv6e/v7/Trainium2/3 | — Kimi K2.7-Code、MiniMax M3、DeepSeekv4、GLM5 - GB200 NVL72 vs MI355X vs B200 vs GB300 NVL72,™ TPUv6e/v7/Trainium2/3 (1,244 stars, Apache-2.0) - [kubeai](https://madewithwhat.net/vllm/project/kubeai/): AI Inference Operator for Kubernetes. The easiest way to serve ML models in production. Supports VLMs, LLMs, embeddings, and speech-to-text. (1,222 stars, Apache-2.0) - [BricksLLM](https://madewithwhat.net/vllm/project/bricksllm/): Enterprise-grade API gateway that helps you monitor and impose cost or rate limits per API key. Get fine-grained access control and monitoring per user, application, or environment. Supports OpenAI, Azure OpenAI, Anthropic, vLLM, and open-source LLMs. (1,217 stars, MIT) - [GPTQModel](https://madewithwhat.net/vllm/project/gptqmodel/): LLM model quantization (compression) toolkit with HW acceleration support for Nvidia, AMD, Intel GPU and Intel/AMD/Apple CPU via HF, vLLM, and SGLang. (1,205 stars) - [prometheus-eval](https://madewithwhat.net/vllm/project/prometheus-eval/): Evaluate your LLM's response with Prometheus and GPT4 (1,102 stars, Apache-2.0) - [kvcached](https://madewithwhat.net/vllm/project/kvcached/): Virtualized Elastic KV Cache for Dynamic GPU Sharing and Beyond (1,097 stars, Apache-2.0) - [tiny-vllm](https://madewithwhat.net/vllm/project/tiny-vllm/): Build your own high performance LLM inference engine in C++ and CUDA - a smaller version of vLLM (919 stars, Apache-2.0) - [llm_note](https://madewithwhat.net/vllm/project/llm-note/): LLM notes, including model inference, transformer model structure, and llm framework code analysis notes. (883 stars) - [llmcord](https://madewithwhat.net/vllm/project/llmcord/): Make Discord your LLM frontend - Supports any OpenAI compatible API (OpenRouter, Ollama and more) (814 stars, MIT) - [BambooAI](https://madewithwhat.net/vllm/project/bambooai/): A Python library powered by Language Models (LLMs) for conversational data discovery and analysis. (782 stars, MIT) - [local_ai_ocr](https://madewithwhat.net/vllm/project/local-ai-ocr/): An local, offline (after initial setup), portable OCR software that can process images and PDF files, using DeepSeek-OCR-2 AI (running directly on your machine). (773 stars, Apache-2.0) - [vidur](https://madewithwhat.net/vllm/project/vidur/): Accurate, large-scale, and extensible simulator for LLM inference Systems (641 stars, MIT) - [efficientsam3](https://madewithwhat.net/vllm/project/efficientsam3/): EfficientSAM3 compresses SAM3 into lightweight, edge-friendly models via progressive knowledge distillation for fast promptable concept segmentation and tracking. (628 stars) - [AI-Bank-Statement-Document-Automation-By-LLM-And-Personal-Finanical-Analysis-Prediction](https://madewithwhat.net/vllm/project/ai-bank-statement-document-automation-by-llm-and-personal-finanical-analysis-prediction/): AI Bank Statement Document Automation By LLM model and Personal Finanical Analysis (597 stars, Apache-2.0) - [RoboBrain](https://madewithwhat.net/vllm/project/robobrain/): [CVPR 2025] RoboBrain: A Unified Brain Model for Robotic Manipulation from Abstract to Concrete. Official Repository. (558 stars, Apache-2.0) - [crater](https://madewithwhat.net/vllm/project/crater/): Crater is a cloud-native AI training & inference platform. (542 stars, Apache-2.0) - [openinfer](https://madewithwhat.net/vllm/project/openinfer/): Pure Rust + CUDA LLM inference engine — no PyTorch, OpenAI-compatible, serves Qwen3 to Kimi-K2 (535 stars, Apache-2.0) - [restai](https://madewithwhat.net/vllm/project/restai/): RESTai is an AIaaS (AI as a Service) open-source platform. Supports many public and local LLM suported by Ollama/vLLM/etc. Precise embeddings usage, tuning, analytics etc. Built-in image/audio generation with dynamic loading generators. Live chat deployment. Built-in block based graphical language. Prompt versioning and much more... (512 stars, Apache-2.0) - [vllm-playground](https://madewithwhat.net/vllm/project/vllm-playground/): A modern web interface for managing and interacting with vLLM servers (www.github.com/vllm-project/vllm). Supports both GPU and CPU modes, with special optimizations for macOS Apple Silicon and enterprise deployment on OpenShift/Kubernetes. (495 stars, Apache-2.0) - [ome](https://madewithwhat.net/vllm/project/ome/): Open Model Engine (OME) — Kubernetes operator for LLM serving, GPU scheduling, and model lifecycle management. Works with SGLang, vLLM, TensorRT-LLM, and Triton (480 stars, Apache-2.0) - [worker-vllm](https://madewithwhat.net/vllm/project/worker-vllm/): The Runpod worker template for serving our large language model endpoints. Powered by vLLM. (456 stars, MIT) - [KVarN](https://madewithwhat.net/vllm/project/kvarn/): KVarN is a native vLLM KV-cache quantization backend for your agents: 3-5x more context, throughput above FP16, and FP16-level accuracy. Calibration-free, one flag. (437 stars, Apache-2.0) - [DeepSeek-OCR-WebUI](https://madewithwhat.net/vllm/project/deepseek-ocr-webui/): Ready-to-use DeepSeek-OCR Web UI | Modern Interface | 7 Recognition Modes | Batch Processing | Real-time Logging | Fully Responsive (436 stars, MIT) - [taOS](https://madewithwhat.net/vllm/project/taos/): Self-hosted AI agent OS. Your memory, chat, agents, and files stay on hardware you own, offline by default, cloud by choice. Offline AI memory (taOSmd), self-hosted multi-framework group chat, a full web desktop + app store, and auto-clustering across the consumer hardware you already have (Orange/Raspberry Pi, Mac mini, gaming PC). (430 stars, AGPL-3.0) - [nodetool](https://madewithwhat.net/vllm/project/nodetool/): The open creative AI workspace (426 stars, AGPL-3.0) - [Qwen3.6-27B-AEON-Ultimate-Uncensored-DFlash](https://madewithwhat.net/vllm/project/qwen3-6-27b-aeon-ultimate-uncensored-dflash/): Fully uncensored, capability-enhanced abliteration of Qwen3.6-27B. NVFP4 + z-lab DFlash speculative decoding (n=12) on the unified ghcr.io/aeon-7/aeon-vllm-ultimate:latest container, tuned for long-context draft acceptance on DGX Spark. 6 HF variants (BF16/NVFP4/MTP/MTP-XS), docker-compose, and QuickStart. (415 stars, Apache-2.0) - [topicGPT](https://madewithwhat.net/vllm/project/topicgpt/): TopicGPT: A Prompt-Based Framework for Topic Modeling [NAACL'24] (411 stars, MIT) - [OneCompression](https://madewithwhat.net/vllm/project/onecompression/): Python package for LLM compression (397 stars, MIT) - [super-json-mode](https://madewithwhat.net/vllm/project/super-json-mode/): Low latency JSON generation using LLMs (396 stars) - [smg](https://madewithwhat.net/vllm/project/smg/): Engine-agnostic LLM gateway in Rust. Full OpenAI & Anthropic API compatibility across vLLM, TRT-LLM, TokenSpeed, SGLang, OpenAI, Gemini & more. Industry-first gRPC pipeline, KV cache-aware routing, chat history, tokenization caching, Responses API, embeddings, WASM plugins, MCP, and multi-tenant auth. (389 stars, Apache-2.0) - [Lvllm](https://madewithwhat.net/vllm/project/lvllm/): LvLLM is a special NUMA extension of vllm that makes full use of CPU and memory resources, reduces GPU memory requirements, and features an efficient GPU parallel and NUMA parallel architecture, supporting hybrid inference for MOE large models. (384 stars, Apache-2.0) - [sparkrun](https://madewithwhat.net/vllm/project/sparkrun/): sparkrun - launch, manage, and stop LLM inference workloads on NVIDIA DGX Spark systems (383 stars, Apache-2.0) - [TinyLLM](https://madewithwhat.net/vllm/project/tinyllm/): Setup and run a local LLM and Chatbot using consumer grade hardware. (343 stars, MIT) - [llmaz](https://madewithwhat.net/vllm/project/llmaz/): Easy, advanced inference platform for large language models on Kubernetes. Star to support our work! (307 stars, Apache-2.0) - [unified-cache-management](https://madewithwhat.net/vllm/project/unified-cache-management/): Persist and reuse KV Cache to speedup your LLM. (302 stars, MIT) - [flama](https://madewithwhat.net/vllm/project/flama/): The production framework for Predictive and Generative AI. Serve any model as an API in one line, with OpenAI/Anthropic/Ollama-compatible endpoints, a built-in chat UI, and native MCP. (292 stars, Apache-2.0) - [xinfer](https://madewithwhat.net/vllm/project/xinfer/): Blazing-fast LLM inference in pure Rust. No PyTorch and Python runtime. (286 stars, MIT) - [llm-benchmark](https://madewithwhat.net/vllm/project/llm-benchmark/): LLM ,。 (267 stars, MIT) - [SEAgent](https://madewithwhat.net/vllm/project/seagent/): [ICML-2026] Official implementation of "SEAgent: Self-Evolving Computer Use Agent with Autonomous Learning from Experience" (257 stars) - [olla](https://madewithwhat.net/vllm/project/olla/): High-performance lightweight proxy and load balancer for LLM infrastructure. Intelligent routing, automatic failover and unified model discovery across local and remote inference backends. (256 stars, Apache-2.0) - [Namo-R1](https://madewithwhat.net/vllm/project/namo-r1/): A CPU Realtime VLM in 500M. Surpassed Moondream2 and SmolVLM. Training from scratch with ease. (256 stars) - [matrixhub](https://madewithwhat.net/vllm/project/matrixhub/): An Open-source, self-hosted AI model hub with Hugging Face compatibility, accelerating vLLM/SGLang performance. (255 stars, Apache-2.0) - [gpt_server](https://madewithwhat.net/vllm/project/gpt-server/): gpt_serverLLMs、Embedding、Reranker、ASR、TTS、、。 (253 stars, Apache-2.0) - [hass_local_openai_llm](https://madewithwhat.net/vllm/project/hass-local-openai-llm/): Home Assistant LLM integration for local OpenAI-compatible services (llamacpp, vllm, etc) (223 stars, Apache-2.0) - [qwen3.6-windows-server](https://madewithwhat.net/vllm/project/qwen3-6-windows-server/): One-click Qwen3.6-27B inference on Windows. 158 tok/s on RTX 5090, 72 tok/s on RTX 3090. Native, no WSL, no Docker, no telemetry. (222 stars) - [nano-vllm](https://madewithwhat.net/vllm/project/nano-vllm/): a fun and educational take on vLLM (211 stars, Apache-2.0) - [TorchSpec](https://madewithwhat.net/vllm/project/torchspec/): A PyTorch native library for training speculative decoding models (198 stars, MIT) - [LMeterX](https://madewithwhat.net/vllm/project/lmeterx/): A general-purpose API load testing platform that supports LLM services and business HTTP interfaces, enabling one-click performance testing, result comparison, and AI-powered intelligent analysis and summarization. API , HTTP ,、 AI (196 stars, Apache-2.0) - [GutenOCR](https://madewithwhat.net/vllm/project/gutenocr/): Open-source tools for training and evaluating Vision Language Models for OCR (189 stars, Apache-2.0) - [nextjs-vllm-ui](https://madewithwhat.net/vllm/project/nextjs-vllm-ui/): Fully-featured, beautiful web interface for vLLM - built with NextJS. (187 stars, MIT) - [RexSeek](https://madewithwhat.net/vllm/project/rexseek/): [ICCV2025] Referring any person or objects given a natural language description. Code base for RexSeek and HumanRef Benchmark (184 stars) - [pegaflow](https://madewithwhat.net/vllm/project/pegaflow/): High-performance KV cache storage for LLM inference — GPU offloading, SSD caching, and cross-node sharing via RDMA. Works with vLLM and SGLang. (170 stars, Apache-2.0) - [booster](https://madewithwhat.net/vllm/project/booster/): Booster - open accelerator for LLM models. Better inference and debugging for AI hackers (169 stars) - [grps](https://madewithwhat.net/vllm/project/grps/): Deep Learning Deployment Framework: Supports tf/torch/trt/trtllm/vllm and other NN frameworks. Support dynamic batching, and streaming modes. It is dual-language compatible with Python and C++, offering scalability, extensibility, and high performance. It helps users quickly deploy models and provide services through HTTP/RPC interfaces. (168 stars, Apache-2.0) - [SGI-Bench](https://madewithwhat.net/vllm/project/sgi-bench/): Probing Scientific General Intelligence of LLMs with Scientist-Aligned Workflows (167 stars, MIT) - [LLMKube](https://madewithwhat.net/vllm/project/llmkube/): Kubernetes operator for self-hosted LLM inference across a heterogeneous GPU fleet: NVIDIA CUDA, AMD Vulkan, and Apple Silicon Metal. Runtimes: llama.cpp, vLLM, TGI, mlx-server. Multi-GPU sharding, model caching, OpenAI-compatible endpoints. Apache-2.0, run across homelab and on-prem fleets, actively developed. (166 stars, Apache-2.0) - [CacheRoute](https://madewithwhat.net/vllm/project/cacheroute/): CacheRoute is an innovative LLM scheduling scheme dedicated to enabling flexible KV cache reuse across LLM systems, improving task performance and system efficiency. (157 stars, Apache-2.0) - [One-Eval](https://madewithwhat.net/vllm/project/one-eval/): Automated system for LLM evaluation via agents. Doc as below: (153 stars, Apache-2.0) - [cognithor](https://madewithwhat.net/vllm/project/cognithor/): Cognithor · Agent OS: Local-first autonomous agent operating system. 19 LLM providers, 18 channels, 145 MCP tools, 6-tier memory, Agent Packs marketplace, zero telemetry. Python 3.12+, Apache 2.0. (150 stars, Apache-2.0) - [glm-5.2-sm120](https://madewithwhat.net/vllm/project/glm-5-2-sm120/): GLM-5.2-NVFP4-REAP-469B serving on SM120 (4× RTX PRO 6000 Blackwell) — one-command vLLM launch recipe, 250K context, DeepSeek Sparse Attention + MTP speculative decode (147 stars) - [docmind-ai-llm](https://madewithwhat.net/vllm/project/docmind-ai-llm/): DocMind AI is a powerful, open-source Streamlit application leveraging LlamaIndex, LangGraph, and local Large Language Models (LLMs) via Ollama, LMStudio, llama.cpp, or vLLM for advanced document analysis. Analyze, summarize, and extract insights from a wide array of file formats, securely and privately, all offline. (137 stars, MIT) - [llamactl](https://madewithwhat.net/vllm/project/llamactl/): Unified management and routing for llama.cpp, MLX and vLLM models with web dashboard. (135 stars, MIT) - [promptmask](https://madewithwhat.net/vllm/project/promptmask/): Never give AI companies your secrets! A local LLM-based privacy filter for LLM users. Seamless integration with your existing AI tools as a Python library / OpenAI SDK replacement / API Gatetway / Web Server. (134 stars, MIT) - [kube-llmops](https://madewithwhat.net/vllm/project/kube-llmops/): A project worth exploring. (122 stars, Apache-2.0) - [sndr_core_engine](https://madewithwhat.net/vllm/project/sndr-core-engine/): SNDR Core Engine (Genesis) — vLLM runtime patch-overlay for Qwen3.6 + Gemma4 on consumer NVIDIA (Ampere sm_86, 2× A5000/3090). Qwen3.6-35B-A3B FP8 ~240 tok/s, 27B-int4 hybrid GDN+Mamba, Gemma4 26B/31B AWQ, 256K ctx. 321 patches: TurboQuant k8v4 KV, MTP/DFlash spec-decode, FULL cudagraph, hybrid GDN. vLLM pin dev424 + Control Center GUI. (122 stars, Apache-2.0) - [voidllm](https://madewithwhat.net/vllm/project/voidllm/): Privacy-first LLM proxy and AI gateway — load balancing, multi-provider routing, API key management, usage tracking, rate limiting. Self-hosted. Zero knowledge of your prompts. (118 stars) - [photon_infer](https://madewithwhat.net/vllm/project/photon-infer/): A High-Performance LLM Inference Engine with vLLM-Style Continuous Batching (118 stars, MIT) - [ppt2desc](https://madewithwhat.net/vllm/project/ppt2desc/): Convert PowerPoint files into semantically rich text using vision language models (113 stars, MIT) - [vector-inference](https://madewithwhat.net/vllm/project/vector-inference/): Efficient LLM inference on Slurm clusters. (106 stars, Apache-2.0) - [dgx-spark-vllm-setup](https://madewithwhat.net/vllm/project/dgx-spark-vllm-setup/): One-command vLLM installation for NVIDIA DGX Spark with Blackwell GB10 GPUs (sm_121 architecture) (102 stars, MIT) - [llm-inference](https://madewithwhat.net/vllm/project/llm-inference/): llm-inference is a platform for publishing and managing llm inference, providing a wide range of out-of-the-box features for model deployment, such as UI, RESTful API, auto-scaling, computing resource management, monitoring, and more. (95 stars, Apache-2.0) - [airunway](https://madewithwhat.net/vllm/project/airunway/): Kubernetes-native platform for deploying and managing AI inference across multiple providers (95 stars, Apache-2.0) - [llmariner](https://madewithwhat.net/vllm/project/llmariner/): Extensible generative AI platform on Kubernetes with OpenAI-compatible APIs. (95 stars, Apache-2.0) - [FlowSteer](https://madewithwhat.net/vllm/project/flowsteer/): FlowSteer: agents designing agentic workflows via reinforced progressive canvas editing. (94 stars) - [nano-vllm-v1](https://madewithwhat.net/vllm/project/nano-vllm-v1/): Nano vLLM with vLLM v1's request scheduling strategy and chunked prefill (92 stars, MIT) - [Chinese-MedQA-Qwen2](https://madewithwhat.net/vllm/project/chinese-medqa-qwen2/): Qwen2+SFT+DPO, SFTTrainer/DPOTrainer/TRPOTrainer,,(neo4j, milvus, LDA, )。, vllm, vllm API embedder + Reranker RAG 。 MDAgents , vllm api 。 (89 stars) - [LLMOne](https://madewithwhat.net/vllm/project/llmone/): Enterprise-grade LLM automated deployment tool that makes AI servers truly "plug-and-play". (87 stars, MulanPSL-2.0) - [End-to-End-Agentic-Ai-Automation-Lab](https://madewithwhat.net/vllm/project/end-to-end-agentic-ai-automation-lab/): This repository contains hands-on projects, code examples, and deployment workflows. Explore multi-agent systems, LangChain, LangGraph, AutoGen, CrewAI, RAG, MCP, automation with n8n, and scalable agent deployment using Docker, AWS, and BentoML. (85 stars, MIT) - [SciEvalKit](https://madewithwhat.net/vllm/project/scievalkit/): A unified evaluation toolkit and leaderboard for rigorously assessing the scientific intelligence of large language and vision–language models across the full research workflow. (85 stars, Apache-2.0) - [granite-switch](https://madewithwhat.net/vllm/project/granite-switch/): Granite Switch — Build AI models like you build software (84 stars, Apache-2.0) - [club-5060ti](https://madewithwhat.net/vllm/project/club-5060ti/): Practical local LLM recipes and benchmarks for RTX 5060 Ti setups (84 stars, MIT) - [ray_vllm_inference](https://madewithwhat.net/vllm/project/ray-vllm-inference/): A simple service that integrates vLLM with Ray Serve for fast and scalable LLM serving. (79 stars, Apache-2.0) - [SearchAgent-X](https://madewithwhat.net/vllm/project/searchagent-x/): A High-Efficiency System of Large Language Model Based Search Agents (79 stars) - [sample-genai-on-eks-starter-kit](https://madewithwhat.net/vllm/project/sample-genai-on-eks-starter-kit/): A comprehensive toolkit for deploying production-ready Generative AI infrastructure on Amazon EKS. Includes pre-configured components for: AI Gateway (LiteLLM) LLM Serving (vLLM, SGLang, Ollama) Vector Databases, Embedding Models (TEI) Observability (Langfuse, Phoenix) etc. Fast-track your GenAI deployment with Kubernetes (77 stars, MIT-0) - [turboquant](https://madewithwhat.net/vllm/project/turboquant/): First open-source implementation of Google TurboQuant (ICLR 2026) -- near-optimal KV cache compression for LLM inference. 5x compression with near-zero quality loss. (76 stars, MIT) - [julius](https://madewithwhat.net/vllm/project/julius/): Simple LLM service identification - translate IP:Port to Ollama, vLLM, LiteLLM, or 60+ other AI services in seconds (76 stars, Apache-2.0) - [llm-gateway](https://madewithwhat.net/vllm/project/llm-gateway/): Zero trust LLM gateway. OpenAI-compatible proxy with semantic routing and load balancing across OpenAI, Anthropic, Ollama, vLLM, and any compatible backend. Identity-based access, virtual API keys, and end-to-end encryption via OpenZiti (75 stars, Apache-2.0) - [Soup](https://madewithwhat.net/vllm/project/soup/): Soup turns the pain of LLM fine-tuning into a simple workflow. One config, one command, done. (75 stars, Apache-2.0) - [vCache](https://madewithwhat.net/vllm/project/vcache/): Reliable and Efficient Semantic Prompt Caching with vCache (75 stars) - [vllm-factory](https://madewithwhat.net/vllm/project/vllm-factory/): Production inference for encoder models - ColBERT, GLiNER, ColPali, embeddings etc. - as vLLM plugins for online and in-process deployment (73 stars, Apache-2.0) - [Decoding-Tree-Sketching](https://madewithwhat.net/vllm/project/decoding-tree-sketching/): [ICML 2026] Decoding Tree Sketching (DTS): a training-free & model agonistic & plug-in framework for LLM parallel reasoning. (71 stars, MIT) - [heron](https://madewithwhat.net/vllm/project/heron/): Agent and LLM API performance monitoring via network packet probe. Measures performance of OpenClaw, Claude, Codex, DeepAgents and more — deployed on the provider side, no SDK changes required. (67 stars, Apache-2.0) - [MU-GOT](https://madewithwhat.net/vllm/project/mu-got/): PDF Parsing Tool: GOT's vLLM acceleration implementation, MinerU for layout recognition, and GOT for table formula parsing. (66 stars) - [advanced-deep-research](https://madewithwhat.net/vllm/project/advanced-deep-research/): Automated Deep Research with LLMs, web search, paper parsing, and didactic summarization. (66 stars, MIT) - [nano-kvllm](https://madewithwhat.net/vllm/project/nano-kvllm/): This project aims to provide a high effective KV cache manage framework for llm inference and improve memory utilization and inference speed. (66 stars, MIT) - [embodied-agents](https://madewithwhat.net/vllm/project/embodied-agents/): EmbodiedAgents is a fully-loaded ROS2 based framework for creating interactive physical agents that can understand, remember, and act upon contextual information from their environment. (63 stars, MIT) - [LLM-Inference-Bench](https://madewithwhat.net/vllm/project/llm-inference-bench/): LLM-Inference-Bench (62 stars, BSD-3-Clause) - [llm-vscode-inference-server](https://madewithwhat.net/vllm/project/llm-vscode-inference-server/): An endpoint server for efficiently serving quantized open-source LLMs for code. (59 stars, Apache-2.0) - [IntelliChat](https://madewithwhat.net/vllm/project/intellichat/): Modern AI chatbot supporting multiple LLMs. Switch between Gemini, Mistral, Llama, Claude and ChatGPT. (58 stars, MIT) - [LogSentinelAI](https://madewithwhat.net/vllm/project/logsentinelai/): LLM-powered security log analyzer: detect threats & anomalies with zero regex — just declare a Pydantic schema. Real-time Telegram alerts, SIEM-ready with Elasticsearch/Kibana. Supports OpenAI, Ollama, vLLM. (57 stars, MIT) - [stopwatch](https://madewithwhat.net/vllm/project/stopwatch/): A tool for benchmarking LLMs on Modal (56 stars, MIT) - [vllm-router](https://madewithwhat.net/vllm/project/vllm-router/): vLLM Router (55 stars, Apache-2.0) - [maru](https://madewithwhat.net/vllm/project/maru/): High-Performance KV Cache Storage Engine on CXL Shared Memory for LLM Inference (55 stars, Apache-2.0) - [vllm-chatterbox-stream](https://madewithwhat.net/vllm/project/vllm-chatterbox-stream/): OpenAI-compatible multilingual TTS server — Chatterbox on vLLM with real-time PCM audio streaming, low time-to-first-byte (~0.7 s), voice cloning, and 23 languages. (55 stars, MIT) - [parsehawk](https://madewithwhat.net/vllm/project/parsehawk/): Turn documents into structured JSON with local-first document AI. Run 100% locally by default, with API, CLI, and Web UI. (53 stars, Apache-2.0) - [vllm-dflash](https://madewithwhat.net/vllm/project/vllm-dflash/): DFlash vLLM for DGX Spark — Plug & Play Block-Diffusion Speculative Decoding (52 stars, Apache-2.0) - [TD-Pipe](https://madewithwhat.net/vllm/project/td-pipe/): [ICPP'25] TD-Pipe: Temporally-Disaggregated Pipeline Parallelism Architecture for High-Throughput LLM Inference (52 stars, Apache-2.0) - [spark_vllm_docker](https://madewithwhat.net/vllm/project/spark-vllm-docker/): DGX Spark / GB10 vLLM Docker stack for large-model serving, presets, patches, and validation notes. (51 stars, Apache-2.0) - [arks](https://madewithwhat.net/vllm/project/arks/): Arks is a cloud-native inference framework running on Kubernetes (51 stars, Apache-2.0) - [xinity-ai](https://madewithwhat.net/vllm/project/xinity-ai/): The open-source AI platform for enterprises that can't send data to the cloud. OpenAI-compatible API, full management dashboard, zero data egress. (50 stars, Apache-2.0) - [VLM-Batch-Deployment](https://madewithwhat.net/vllm/project/vlm-batch-deployment/): Batch Deployment for Document Parsing with AWS Batch & Qwen-2.5-VL (49 stars) - [blackbird](https://madewithwhat.net/vllm/project/blackbird/): A high-performance RDMA distributed file system for fast LLM Inference and GPU Training. (48 stars, Apache-2.0) - [turboquant-vllm](https://madewithwhat.net/vllm/project/turboquant-vllm/): TurboQuant KV cache compression plugin for vLLM — asymmetric K/V, 8 models validated, consumer GPUs (48 stars, Apache-2.0) - [vllm-awq4-qwen](https://madewithwhat.net/vllm/project/vllm-awq4-qwen/): vLLM Qwen 3.6-27B (AWQ-INT4) + DFlash speculative decoding on AMD Strix Halo (gfx1151 iGPU, 128 GB UMA, ROCm 7.13). 24.8 t/s single-stream, vision, tool calling, 256K context, OpenAI-compatible, Docker. Matches DGX Spark FP8+DFlash+MTP at a third of the cost. No CUDA. (47 stars, Unlicense) - [localvoxtral](https://madewithwhat.net/vllm/project/localvoxtral/): Native macOS menu bar app for realtime dictation with optional LLM polishing. Connects to any OpenAI Realtime-compatible backend — fully local on Apple Silicon with voxmlx + mlx-swift-lm. (45 stars, MIT) - [ElasticMM](https://madewithwhat.net/vllm/project/elasticmm/): ElasticMM: Elastic and Efficient MLLM Serving System (44 stars) - [bench-loop](https://madewithwhat.net/vllm/project/bench-loop/): Local-first CLI for benchmarking LLMs on real hardware — quality, speed, reliability, and a real multi-turn agent loop. (39 stars, MIT) - [AGmind](https://madewithwhat.net/vllm/project/agmind/): Private LLM/RAG platform in one command for NVIDIA DGX Spark / GB10 (arm64). Validated on real hardware. (38 stars) - [Climatik-Project](https://madewithwhat.net/vllm/project/climatik-project/): Carbon Limiting Auto Tuning for Kubernetes (38 stars, Apache-2.0) - [modelship](https://madewithwhat.net/vllm/project/modelship/): Self-hosted, OpenAI-compatible inference for the agentic era: reasoning LLMs, universal tool calling, and the Responses API alongside embeddings, speech, and image models — many models sharing your GPUs, one gateway. Powered by Ray Serve. (37 stars, Apache-2.0) - [Deskdrop](https://madewithwhat.net/vllm/project/deskdrop/): Android keyboard with local AI (Ollama, Whisper, MCP) or cloud (Gemini, Groq, OpenAI) (36 stars, GPL-3.0) - [ccl](https://madewithwhat.net/vllm/project/ccl/): Hit your limit? Need privacy? Just swap the model, everything else stays (36 stars, MIT) - [AIfred-Intelligence](https://madewithwhat.net/vllm/project/aifred-intelligence/): AIfred-Intelligence — self-hosted Multi-Agent Assistant with Debate Modes (Symposion/Tribunal), Voice (STT + Streaming-TTS), RAG with Long-Term Memory, Web Research and Tool-Calling. Reachable via Web-UI or Telegram/Discord/Email/EPIM. Multi-Backend: llama-swap, Ollama, vLLM, TabbyAPI. (33 stars) - [vllm-gb10](https://madewithwhat.net/vllm/project/vllm-gb10/): Bleeding edge vLLM Docker image for the NVIDIA DGX Spark (GB10 / sm_121a). (31 stars, MIT) - [go-llm-proxy](https://madewithwhat.net/vllm/project/go-llm-proxy/): Lightweight proxy for LLM (31 stars, MIT) - [imp](https://madewithwhat.net/vllm/project/imp/): From-scratch C++/CUDA inference engine for the NVIDIA RTX 5090 (sm_120a) — the best single-GPU backend for agentic AI: tool calling, long-context loops, reasoning and concurrent sub-agents on top of the fastest single-stream decode on the 5090 (beats llama.cpp, at-or-ahead of vLLM on NVFP4). 100% written by Claude Code. (30 stars, MIT) - [Local-CLI](https://madewithwhat.net/vllm/project/local-cli/): AI Coding Assistant CLI for offline enterprise environments - Local LLM platform with Plan & Execute architecture, Supervised Mode, and auto-update system (30 stars, MIT) - [openmagpie](https://madewithwhat.net/vllm/project/openmagpie/): Open-source, self-hostable social listening. (29 stars) - [happy_vllm](https://madewithwhat.net/vllm/project/happy-vllm/): A REST API for vLLM, production ready (29 stars, AGPL-3.0) - [PyLangPipe](https://madewithwhat.net/vllm/project/pylangpipe/): a simple lightweight large language model pipeline framework. (28 stars, Apache-2.0) - [LLM-fine-tuner](https://madewithwhat.net/vllm/project/llm-fine-tuner/): Powerful no-code LLM fine-tuner: upload data → train → deploy in minutes. Unsloth 2-5× acceleration · QLoRA/DPO/RLHF/PPO/ORPO · Reward Model training · GGUF export · vLLM inference · BLEU/ROUGE/BERTScore · full CLI · Heretic Mode to unlock full model potential (28 stars, GPL-3.0) - [Immune](https://madewithwhat.net/vllm/project/immune/): [CVPR2025] Official Repository for IMMUNE: Improving Safety Against Jailbreaks in Multi-modal LLMs via Inference-Time Alignment (28 stars) - [vllm-windows-build](https://madewithwhat.net/vllm/project/vllm-windows-build/): Native Windows vLLM 0.24.0: Python 3.13, CUDA 12.8, PyTorch 2.11 cu128, RTX 30/40/50 wheel, hardened portable installer with legacy PowerShell support, Qwen3-VL/FlashAttention fixes, Rust frontend/tool parser, OpenAI server fixes, and 10 KV-cache compression dtypes. (28 stars, MIT) - [fastassert](https://madewithwhat.net/vllm/project/fastassert/): Dockerized LLM inference server with constrained output (JSON mode), built on top of vLLM and outlines. Faster, cheaper and without rate limits. Compare the quality and latency to your current LLM API provider. (27 stars) - [conch](https://madewithwhat.net/vllm/project/conch/): A "standard library" of Triton kernels. (26 stars, Apache-2.0) - [huggingface-estimate](https://madewithwhat.net/vllm/project/huggingface-estimate/): A web-based memory usage and performance calculator for Huggingface GGUF models (26 stars) - [mixinputs](https://madewithwhat.net/vllm/project/mixinputs/): Official implementation for Text Generation Beyond Discrete Token Sampling (26 stars, Apache-2.0) - [vllm-windows](https://madewithwhat.net/vllm/project/vllm-windows/): Patched native-Windows build of vLLM. Three Windows-specific fixes (CPU-relay for Gloo, Qwen3 reasoning parser, wildcard model name) on top of SystemPanic 0.19.0. Engine for devnen/qwen3.6-windows-server. (26 stars, Apache-2.0) - [tandemn-tuna](https://madewithwhat.net/vllm/project/tandemn-tuna/): A hybrid router that uses Spot GPU instances to reduce costs and Serverless GPUs for making Cold Starts faster. (26 stars, MIT) - [orpheus-streaming](https://madewithwhat.net/vllm/project/orpheus-streaming/): Orpheus TTS Server with streaming support (TTFB ~160ms) (26 stars, Apache-2.0) - [avp-python](https://madewithwhat.net/vllm/project/avp-python/): Python SDK for Agent Vector Protocol – transfer KV-cache between LLM agents instead of text (25 stars, Apache-2.0) - [nexusquant](https://madewithwhat.net/vllm/project/nexusquant/): Training-free KV cache compression via E8 lattice VQ. 2-bit KV that preserves retrieval (30/30 NIAH vs TurboQuant 0/30). Calibration-free, 9 architectures validated. (25 stars) - [multi-turboquant](https://madewithwhat.net/vllm/project/multi-turboquant/): Unified KV cache compression for LLM inference — TurboQuant, IsoQuant, PlanarQuant, TriAttention. 10 methods, GPU-validated, multi-GPU planner. Compress KV cache 5-80x to run bigger models, longer context, more agents on your GPU. (24 stars, MIT) - [LLM-VoIP-Caller](https://madewithwhat.net/vllm/project/llm-voip-caller/): This project is the backend engine for a fully autonomous AI-powered call center. It integrates a large language model (LLM), speech recognition, and text-to-speech to manage real-time phone conversations via Asterisk. (24 stars) - [ChaosEngineAI](https://madewithwhat.net/vllm/project/chaosengineai/): Local AI workstation — discover, run, chat, benchmark, and generate images from open-weight models. DFlash/DDTree speculative decoding, TurboQuant & TriAttention cache compression strategies, MLX + llama.cpp + vLLM + MTPLX backends. (24 stars, Apache-2.0) - [macbench](https://madewithwhat.net/vllm/project/macbench/): Probing the limitations of multimodal language models for chemistry and materials research (24 stars, MIT) - [lm-fly](https://madewithwhat.net/vllm/project/lm-fly/): , LLM (24 stars, MIT) - [llm-swarm-router](https://madewithwhat.net/vllm/project/llm-swarm-router/): Run the LLM Swarm Router on machines to distribute Local Ai to the Swarm - More Machines - MORE SPEED (23 stars, MIT) - [llm-bench-rig](https://madewithwhat.net/vllm/project/llm-bench-rig/): Dual-engine (llama.cpp + vLLM) LLM benchmarking pipeline for GGUF & safetensors on NVIDIA GPUs — speed, quality, live dashboard, publishable cards. (23 stars) - [open-agent-sdk-rust](https://madewithwhat.net/vllm/project/open-agent-sdk-rust/): Rust SDK for building AI agents with local OpenAI-compatible servers (LMStudio, Ollama, llama.cpp, vLLM). Features streaming, tools, hooks, retry logic, and comprehensive examples. (23 stars, MIT) - [ai-natural-language-tests](https://madewithwhat.net/vllm/project/ai-natural-language-tests/): Enterprise-grade platform to generate and execute Cypress, Playwright, WebdriverIO, and Appium end-to-end tests from natural language requirements. (Appium is experimental and requires external mobile infrastructure.) (22 stars, AGPL-3.0) - [Comfy-Cozy](https://madewithwhat.net/vllm/project/comfy-cozy/): Say 'dreamier' and your ComfyUI workflow shifts — instantly, reversibly. An AI co-pilot for VFX artists: zero-LLM recipes (dreamier/sharper/faster), 133 MCP tools, EXR-aware vision, workflow.lock provenance, full undo, native sidebar, 5 swappable brains (Claude, GPT, Gemini, Ollama, Nemotron). (22 stars, MIT) - [agentic-loop-engineering-course](https://madewithwhat.net/vllm/project/agentic-loop-engineering-course/): An 18 notebook course that isolates and measures each component of agentic loop engineering on real, industry standard software datasets. (21 stars, MIT) - [Terradev](https://madewithwhat.net/vllm/project/terradev/): An imperative command-line-interface for AI workload orchestration (21 stars, Apache-2.0) - [vllm-surya-ocr](https://madewithwhat.net/vllm/project/vllm-surya-ocr/): OpenAI-compatible, vLLM-served OCR API for the Surya-OCR-2 model — multilingual document OCR (layout + text recognition) with request batching, a local CLI, and Docker packaging. (21 stars, MIT) - [CoCoA](https://madewithwhat.net/vllm/project/cocoa/): [ACL 2026] CoCoA: Collaborative Chain-of-Agents for Parametric-Retrieved Knowledge Synergy (21 stars) - [Aeon-Bench-Pod](https://madewithwhat.net/vllm/project/aeon-bench-pod/): Run the AEON Bench suite on your own hardware: verified HuggingFace pull → serve → benchmark (text · agentic ×3 harnesses · vision · audio · arena · perf) → ed25519-signed attested submit. (20 stars, MIT) - [blackwell-geforce-nvfp4-gemm](https://madewithwhat.net/vllm/project/blackwell-geforce-nvfp4-gemm/): NVFP4 inference on Blackwell GeForce (RTX 5090/5080/5070 Ti/RTX PRO 6000) — SM120 patches for vLLM + FlashInfer + CUTLASS. 175 tok/s on Qwen3.6-35B MoE. (20 stars) - [transcria](https://madewithwhat.net/vllm/project/transcria/): Self-hosted meeting transcription portal — speech-to-text, speaker diarization, LLM-corrected transcripts, structured summaries and Word minutes, on your own GPUs. Flask + PostgreSQL, GDPR audit trail, distributed GPU topologies, docker (19 stars, Apache-2.0) - [Efficient-LLM-Inference-Serving-Systems](https://madewithwhat.net/vllm/project/efficient-llm-inference-serving-systems/): Why is LLM inference slow — and how do you make it fast? A hands-on, first-principles course: roofline → KV cache → quantization → parallelism → vLLM/SGLang, with GPU labs on open models. (19 stars, MIT) - [axetract](https://madewithwhat.net/vllm/project/axetract/): Low-Cost Cross-Domain Web Structured Information Extraction using specialized LoRA adapters. (19 stars, MIT) - [EduRAG-NetworkAssistant](https://madewithwhat.net/vllm/project/edurag-networkassistant/): RAG using LlamaIndex:Computer Network Q&A System powered by LlamaIndex | LlamaIndex - HyDE+ + vLLM +Ragas (19 stars) - [guidance-for-scalable-model-inference-and-agentic-ai-on-amazon-eks](https://madewithwhat.net/vllm/project/guidance-for-scalable-model-inference-and-agentic-ai-on-amazon-eks/): Comprehensive, scalable ML inference architecture using Amazon EKS, leveraging Graviton processors for cost-effective CPU-based inference and GPU instances for accelerated inference. Guidance provides a complete end-to-end platform for deploying LLMs with agentic AI capabilities, including RAG and MCP (19 stars, MIT-0) - [ICE-PIXIU](https://madewithwhat.net/vllm/project/ice-pixiu/): ICE-PIXIU:A Cross-Language Financial Megamodeling Framework (19 stars, Apache-2.0) - [llmq](https://madewithwhat.net/vllm/project/llmq/): A Scheduler for Batched LLM Inference (19 stars) - [mvllm](https://madewithwhat.net/vllm/project/mvllm/): Intelligent load balancer for distributed vLLM server clusters vLLM (19 stars, MIT) - [oxRL](https://madewithwhat.net/vllm/project/oxrl/): A lightweight post-training framework for LLMs and VLMs. 51 algorithms, 38 verified models. Scales with DeepSpeed, vLLM, and Ray. (19 stars, Apache-2.0) - [DAR](https://madewithwhat.net/vllm/project/dar/): Source code for the paper: Hear Both Sides: Efficient Multi-Agent Debate via Diversity-Aware Message Retention (18 stars, MIT) - [rag-colls](https://madewithwhat.net/vllm/project/rag-colls/): Collection of recent advanced RAG techniques. (18 stars, MIT) - [wombatkv](https://madewithwhat.net/vllm/project/wombatkv/): Object-storage-native KV cache for LLM inference & RL. Cross-restart, cross-conversation, cross-engine via shared S3 bucket. (18 stars, Apache-2.0) - [JSE](https://madewithwhat.net/vllm/project/jse/): Local-first desktop assistant for the whole job hunt — scrape listings, match with a local LLM, generate applications in your own voice, and track everything on a Kanban board. (18 stars) - [attach-gateway](https://madewithwhat.net/vllm/project/attach-gateway/): Drop-in OIDC & Google A2A auth + Weaviate memory for Ollama, vLLM and any local LLM server. (18 stars, MIT) - [layered-prefill](https://madewithwhat.net/vllm/project/layered-prefill/): Layered prefill changes the scheduling axis from tokens to layers and removes redundant MoE weight reloads while keeping decode stall free. The result is lower TTFT, lower end-to-end latency, and lower energy per token without hurting TBT stability. (18 stars, MIT) - [Manalyzer](https://madewithwhat.net/vllm/project/manalyzer/): Manalyzer: End-to-end Automated Meta-analysis with Multi-agent System (18 stars) - [dynamic-batching](https://madewithwhat.net/vllm/project/dynamic-batching/): The official repo for the paper "Optimizing LLM Inference Throughput via Memory-aware and SLA-constrained Dynamic Batching" (18 stars, Apache-2.0) - [Aris-AI-Model-Server](https://madewithwhat.net/vllm/project/aris-ai-model-server/): An OpenAI Compatible API which integrates LLM, Embedding and Reranker. LLM、Embedding Reranker OpenAI API (18 stars, Apache-2.0) - [konash](https://madewithwhat.net/vllm/project/konash/): KONASH: Train knowledge agents that search, retrieve, and reason. Based on KARL (Databricks, 2026). (17 stars) - [vllm-doctor](https://madewithwhat.net/vllm/project/vllm-doctor/): Diagnose vLLM inference servers (17 stars, Apache-2.0) - [aixcl](https://madewithwhat.net/vllm/project/aixcl/): Agentic SDLC with local LLM's. (17 stars, Apache-2.0) - [localharness](https://madewithwhat.net/vllm/project/localharness/): An open-source, model-agnostic agent harness for local LLMs. Define agents in YAML (tools, memory, deny-first permissions) and run them against any OpenAI-compatible endpoint: vLLM, Ollama, LM Studio, or llama.cpp. (17 stars, MIT) - [aiburstcloud](https://madewithwhat.net/vllm/project/aiburstcloud/): Dual-mode cloud burst LLM router. Edge-first or cloud-first inference with data sovereignty, cost controls, and automatic failover. https://aiburstcloud.com (17 stars, MIT) - [GLM-spark](https://madewithwhat.net/vllm/project/glm-spark/): Serve GLM-5.2 469B (REAP-pruned, NVFP4) across 3× NVIDIA DGX Spark with vLLM pipeline parallelism — 256K context, production-ready config and patches (17 stars, MIT) - [doc2md](https://madewithwhat.net/vllm/project/doc2md/): Convert pdf and image files into markdown (16 stars, MIT) - [Multi-Agent-Benchmark-Tool](https://madewithwhat.net/vllm/project/multi-agent-benchmark-tool/): Interested in running a conspiracy of agents (technical term) on your Local AI Infra? Who isn't! (16 stars, MIT) - [dual-rtx-6000-blackwell-qwen3.6-27b-fp8](https://madewithwhat.net/vllm/project/dual-rtx-6000-blackwell-qwen3-6-27b-fp8/): Optimized vLLM setup for Qwen3.6-27B-FP8 on dual RTX PRO 6000 Blackwell (192 GB GDDR7, no NVLink); config, benchmark sweep results, and custom chat template with thinking mode off by default. (15 stars, MIT) - [demodel](https://madewithwhat.net/vllm/project/demodel/): Easily boost the speed of pulling your models and datasets from various of inference runtimes. (e.g. HuggingFace, Ollama, vLLM, and more!) (15 stars, MIT) - [vllm-qwen](https://madewithwhat.net/vllm/project/vllm-qwen/): vLLM + Qwen3.6-27B (BF16) OpenAI-compatible inference server on AMD Strix Halo (Ryzen AI Max+ 395, gfx1151). Vision input, 256K context, /v1/responses with separated reasoning, via TheRock ROCm. (15 stars, Unlicense) - [RAG-LCC](https://madewithwhat.net/vllm/project/rag-lcc/): A hands‑on RAG experimentation lab. Largely configurable with debug insights. Classification‑driven corpus construction, filter chains, document loading, chat interaction, Open WebUI integration. Experimental by design and not production‑ready. (15 stars, MIT) - [TokenPowerBench](https://madewithwhat.net/vllm/project/tokenpowerbench/): chenxuniu/TokenSpark-Benchmark-Benchmarking-Power-Consumption-of-LLM-Inference-on-Multi-Node-Clusters (14 stars) - [Batch_LLM_Inference_with_Ray_Data_LLM](https://madewithwhat.net/vllm/project/batch-llm-inference-with-ray-data-llm/): Batch LLM Inference with Ray Data LLM: From Simple to Advanced (13 stars, MIT) - [ainode](https://madewithwhat.net/vllm/project/ainode/): Turn any NVIDIA GPU into a local AI platform. Inference + fine-tuning in your browser. One command to start, automatic clustering. (13 stars, Apache-2.0) - [llm-lab](https://madewithwhat.net/vllm/project/llm-lab/): LLM, Fine Tuning, Llama 2, Gemma, Mixtral, vLLM, LangChain, RAG, ChromaDB, FAISS (13 stars) - [flash-head](https://madewithwhat.net/vllm/project/flash-head/): FlashHead: Efficient Drop-In Replacement for the Classification Head in Language Model Inference (13 stars) - [exocomp](https://madewithwhat.net/vllm/project/exocomp/): Exocomp Agentic Environment for Go (13 stars) - [piqc](https://madewithwhat.net/vllm/project/piqc/): Kubernetes scanner that discovers LLMs running on vLLM and extracts their deployment and runtime facts. (13 stars) - [gpumod](https://madewithwhat.net/vllm/project/gpumod/): GPU Service Manager for LLM workloads on Linux/NVIDIA systems. (13 stars, Apache-2.0) - [SAM](https://madewithwhat.net/vllm/project/sam/): SAM — Smart Agentic Model: CLI coding agent for open-source LLMs. pip install sam-agent (13 stars, MIT) - [trajectorykit](https://madewithwhat.net/vllm/project/trajectorykit/): A lean, local AI research agent framework that runs on your machine (12 stars, MIT) - [periscope](https://madewithwhat.net/vllm/project/periscope/): LLM Performance Testing | K6 + Grafana + InfluxDB | A tiny toolkit for load testing and benchmarking OpenAI-like inference endpoints using K6 + Grafana + InfluxDB (12 stars) - [flexbench](https://madewithwhat.net/vllm/project/flexbench/): Benchmark OpenAI-compatible AI endpoints and AI Accelerators in a reproducible structured way (12 stars, Apache-2.0) - [distributed-inference-vllm](https://madewithwhat.net/vllm/project/distributed-inference-vllm/): Distributed Inference with vLLM (12 stars, Apache-2.0) - [gfx906-fa-vllm](https://madewithwhat.net/vllm/project/gfx906-fa-vllm/): FlashAttention-style custom attention backend for vLLM on AMD MI50/MI60/Radeon VII (gfx906). Downstream fork of mixa3607/ML-gfx906 with replacement HIP kernels and a vllm.general_plugins entry point. (12 stars, Apache-2.0) - [rl-infra-notes](https://madewithwhat.net/vllm/project/rl-infra-notes/): Source-code level analysis of LLM RL training infra: async RL, weight sync, FP8, MoE routing | LLM RL (12 stars) - [cascade](https://madewithwhat.net/vllm/project/cascade/): Extend LLM context windows beyond GPU memory limits with disk-backed KV cache. (11 stars, Apache-2.0) - [vast-coding-llm](https://madewithwhat.net/vllm/project/vast-coding-llm/): Terraform setup for deploying a private coding LLM on Vast.ai with vLLM, Qwen3 Coder, and OpenCode. (11 stars) - [WhiteLotus](https://madewithwhat.net/vllm/project/whitelotus/): LLM inference engine built from scratch in C++. No PyTorch, no frameworks. (11 stars, MIT) - [ask-poddy](https://madewithwhat.net/vllm/project/ask-poddy/): Ask Poddy: Run Open Source LLMs and Embeddings as OpenAI-Compatible Serverless Endpoints (Tutorial) (11 stars, AGPL-3.0) - [agentsculptor](https://madewithwhat.net/vllm/project/agentsculptor/): agentsculptor is an experimental AI-powered development agent designed to analyze, refactor, and extend Python projects automatically. It uses an OpenAI-like planner–executor loop on top of a vLLM backend, combining project context analysis, structured tool calls, and iterative refinement. It has only been tested with gpt-oss-120b via vLLM. (11 stars, Apache-2.0) - [AIMA](https://madewithwhat.net/vllm/project/aima/): AI-Inference-Managed-by-AI: Go binary for managing AI inference on edge devices (11 stars, Apache-2.0) - [dgx_spark](https://madewithwhat.net/vllm/project/dgx-spark/): Multi-model LLM serving for NVIDIA DGX Spark with vLLM, web UI, and tool calling (11 stars, Apache-2.0) - [llmhop](https://madewithwhat.net/vllm/project/llmhop/): Tiny, stateless Go router that dispatches OpenAI-compatible requests to single-model vLLM and sglang backends with zero external dependencies (11 stars, MIT) - [honeycomb-lab](https://madewithwhat.net/vllm/project/honeycomb-lab/): Honeycomb Lab — hex map + OpenAI gateway control plane for a home AI fleet (10 stars, MIT) - [tinycode](https://madewithwhat.net/vllm/project/tinycode/): A slim, local-LLM-first AI coding assistant. TUI, Web UI, and desktop app. Runs air-gapped with Ollama, vLLM, or any OpenAI-compatible endpoint. (10 stars, MIT) - [OdinsList](https://madewithwhat.net/vllm/project/odinslist/): Automated comic cataloging tool that identifies issues directly from cover images using a vision-language model, then cross-references results with the Grand Comics Database and the ComicVine API to generate structured, high-confidence collection data with minimal manual entry. (10 stars, MIT) - [NLP-Reading-List](https://madewithwhat.net/vllm/project/nlp-reading-list/): A curated collection of NLP and LLM resources. Covers essential papers and blogs on Transformers, Reinforcement Learning (RLHF, DPO, GRPO), Mechanistic Interpretability, Scaling Laws, and MLSys. (10 stars) - [openllm-web](https://madewithwhat.net/vllm/project/openllm-web/): Open-source, self-hosted LLM chat application, featuring local-first data storage and real-time streaming responses. (10 stars, MIT) - [AI_Secretary_System](https://madewithwhat.net/vllm/project/ai-secretary-system/): AI-,. XTTS v2, real-time (Vosk/Whisper) offline LLM. - (Vue 3), Telegram-,, fine-tuning pipeline. Self-hosted,,. (10 stars, MIT) - [t2yLLM](https://madewithwhat.net/vllm/project/t2yllm/): A voice assistant with local LLM as a backend (10 stars, MIT) - [langertha](https://madewithwhat.net/vllm/project/langertha/): Perl Framework for AI - Langertha - the viking of AI (10 stars) - [mirage-bench](https://madewithwhat.net/vllm/project/mirage-bench/): Repository for Multililngual Generation, RAG evaluations, and surrogate judge training for Arena RAG leaderboard (NAACL'25) (10 stars, Apache-2.0) ## DevTools (35) - [UniRL](https://madewithwhat.net/vllm/project/unirl/): UniRL is a Framework for Unified Multimodal Model Reinforcement Learning (810 stars) - [LightCompress](https://madewithwhat.net/vllm/project/lightcompress/): [EMNLP 2024 & AAAI 2026] A powerful toolkit for compressing large models including LLMs, VLMs, and video generative models. (733 stars, Apache-2.0) - [FlashTTS](https://madewithwhat.net/vllm/project/flashtts/): SparkTTS、OrpheusTTS,。 (612 stars) - [vllm-cli](https://madewithwhat.net/vllm/project/vllm-cli/): A command-line interface tool for serving LLM using vLLM. (505 stars, MIT) - [apex](https://madewithwhat.net/vllm/project/apex/): AI-powered offensive security testing using autonomous agents, directly in your terminal. (295 stars, Apache-2.0) - [ML-gfx906](https://madewithwhat.net/vllm/project/ml-gfx906/): ML software (llama.cpp, ComfyUI, vLLM) builds for AMD gfx906 GPUs, e.g. Radeon VII / MI50 / MI60 (274 stars, MIT) - [DGX_Spark_Qwen3.5-122B-A10B-AR-INT4](https://madewithwhat.net/vllm/project/dgx-spark-qwen3-5-122b-a10b-ar-int4/): Qwen3.5-122B-A10B on DGX Spark: 28.3 → 51 tok/s (+80%) (274 stars, Apache-2.0) - [vLLM-5090](https://madewithwhat.net/vllm/project/vllm-5090/): vLLM-5090: Docker Container for RTX 5090 + OpenCode (137 stars) - [keda-gpu-scaler](https://madewithwhat.net/vllm/project/keda-gpu-scaler/): KEDA External gRPC Scaler for GPU workloads - native NVML metrics via DaemonSet, no Prometheus required (107 stars, Apache-2.0) - [spark-doctor](https://madewithwhat.net/vllm/project/spark-doctor/): Local diagnostic CLI for NVIDIA DGX Spark (GB10). Detects power caps, UMA pressure, thermal risk, CUDA 13/SM_121 wheel mismatches, Docker issues, and validates vLLM/Ollama/llama.cpp/SGLang recipes. (89 stars, MIT) - [sardeenz](https://madewithwhat.net/vllm/project/sardeenz/): Sardeenz is a proof-of-concept application that allows you to load more than one model on a given GPU. It allows you to add more and more models onto a GPU, until it is fully utilized. (60 stars, Apache-2.0) - [profile](https://madewithwhat.net/vllm/project/profile/): A physics-grounded, cost-aware optimizer for vLLM. (48 stars) - [stuff](https://madewithwhat.net/vllm/project/stuff/): useful Gentoo overlay Curated ebuilds, AI, tools & science (42 stars, GPL-2.0) - [self-expansion](https://madewithwhat.net/vllm/project/self-expansion/): Code for building self-expanding knowledge graphs with Outlines, vLLM, neo4j, and Modal. (37 stars, MIT) - [Deploying-Llama-3.3-70B](https://madewithwhat.net/vllm/project/deploying-llama-3-3-70b/): Serve Llama 3.3 70B (with AWQ quantization) using vLLM and deploy it on BentoCloud. (32 stars, Apache-2.0) - [InferenceBench](https://madewithwhat.net/vllm/project/inferencebench/): Benchmarking Open-Ended Inference Optimization by AI Agents (32 stars, Apache-2.0) - [vllm-cn](https://madewithwhat.net/vllm/project/vllm-cn/): vllm (31 stars) - [hermes](https://madewithwhat.net/vllm/project/hermes/): Policy-driven seamless lazy loading (29 stars, Apache-2.0) - [deepseek-v3-r1-deploy-and-benchmarks](https://madewithwhat.net/vllm/project/deepseek-v3-r1-deploy-and-benchmarks/): DeepSeek-V3, R1 671B on 8xH100 Throughput Benchmarks (22 stars) - [glm-5.2-gb10](https://madewithwhat.net/vllm/project/glm-5-2-gb10/): GLM-5.2 (744B/40B MoE) on a 4× DGX Spark / GB10 (sm_121) cluster: portable Triton sparse-MLA kernels, a data-free expert prune, MTP draft, and a one-script bootstrap. (21 stars, Apache-2.0) - [VisEdit](https://madewithwhat.net/vllm/project/visedit/): [AAAI 2025 oral] Attribution Analysis Meets Model Editing: Advancing Knowledge Correction in Vision Language Models with VisEdit (19 stars) - [DGX-SPARK](https://madewithwhat.net/vllm/project/dgx-spark/): DGX Spark research and tests - containers, benchmarks, and investigation notes for running models on GB10 (SM 12.1) (18 stars) - [llmtop](https://madewithwhat.net/vllm/project/llmtop/): htop for your LLM inference cluster (17 stars) - [IndustryEQA](https://madewithwhat.net/vllm/project/industryeqa/): The official implementation of NeurlPS 2025 D&B paper: IndustryEQA: Pushing the frontiers of Embodied Question Answering in Industrial Scenarios. (16 stars) - [gemma4-12b-vllm-sm120](https://madewithwhat.net/vllm/project/gemma4-12b-vllm-sm120/): Reproducible recipe: serve abliterated Gemma-4-12B (gemma4_unified) at 50-118 tok/s on no-NVLink Blackwell (SM120) via vLLM nightly + ModelOpt FP8/NVFP4 + MTP spec-decode. (15 stars, Apache-2.0) - [vllm-pascal](https://madewithwhat.net/vllm/project/vllm-pascal/): A project worth exploring. (14 stars, MIT) - [EasyLLM](https://madewithwhat.net/vllm/project/easyllm/): Running Large Language Model easily. (13 stars, Apache-2.0) - [vllm_benchmark_block_fp8](https://madewithwhat.net/vllm/project/vllm-benchmark-block-fp8/): Automated Triton w8a8 block FP8 kernel tuning tool for vLLM. Auto-detects model architecture, supports Qwen3-Coder-30B-A3B-Instruct-FP8/DeepSeek-V3/custom models, multi-GPU parallel tuning, and generates optimized kernel configs for quantization. (13 stars) - [vllm-qwen2.5-coder-tool-parser](https://madewithwhat.net/vllm/project/vllm-qwen2-5-coder-tool-parser/): vLLM tool parser for Qwen2.5-Coder models using tag format. (12 stars) - [ai-factory-ops-lab](https://madewithwhat.net/vllm/project/ai-factory-ops-lab/): Hands-on GPU/HPC infrastructure operations: K8s GPU scheduling, HAMi sharing, Slurm, observability & vLLM inference. Learn it free on a laptop; validate on one cheap GPU. (12 stars, MIT) - [down-craft](https://madewithwhat.net/vllm/project/down-craft/): npm pacakge to Craft files into Markdown with ease (11 stars, Apache-2.0) - [nm-vllm-certs](https://madewithwhat.net/vllm/project/nm-vllm-certs/): General Information, model certifications, and benchmarks for nm-vllm enterprise distributions (11 stars) - [WiCo](https://madewithwhat.net/vllm/project/wico/): The official implementation of CVPR Workshop 2025 paper: Window Token Concatenation for Efficient Visual Large Language Models. (11 stars) - [vllm-openwebui-nginx-compose](https://madewithwhat.net/vllm/project/vllm-openwebui-nginx-compose/): Docker Compose setup that integrates vLLM, Open WebUI and NGINX (for SSL termination). (10 stars, MIT) - [spark-pulse](https://madewithwhat.net/vllm/project/spark-pulse/): Spark Pulse is a web control plane for spark-vllm-docker (10 stars) ## Docs (4) - [verl-omni](https://madewithwhat.net/vllm/project/verl-omni/): Multimodal RL training framework for diffusion & omni models (548 stars, Apache-2.0) - [Flash-RL](https://madewithwhat.net/vllm/project/flash-rl/): Implementation for FP8/INT8 Rollout for RL training without performence drop. (305 stars, MIT) - [vllm-cn](https://madewithwhat.net/vllm/project/vllm-cn/): vLLM Documentation in Chinese Simplified / vLLM (185 stars, MIT) - [synthvision](https://madewithwhat.net/vllm/project/synthvision/): Synthetic medical VQA pipeline: 119K images annotated by frontier VLMs, cross-validated at 93% agreement, fine-tuned on 3 model families (2-3B params) (39 stars) ## Dashboards (3) - [spark-dashboard](https://madewithwhat.net/vllm/project/spark-dashboard/): Real-time hardware and LLM inference monitoring — GPU, CPU, memory, and vLLM metrics streamed to a dashboard. (75 stars, MIT) - [vllm-tuner](https://madewithwhat.net/vllm/project/vllm-tuner/): An intelligent tuner for vLLM that automatically monitors GPU metrics, uses Bayesian optimization to tune parameters (65 stars, MIT) - [dspark-vllm-gx10](https://madewithwhat.net/vllm/project/dspark-vllm-gx10/): Two-node DGX Spark/ASUS GX10 DeepSeek V4 Flash DSpark NVFP4 port for vLLM 0.25, with live dashboard and reproducible deployment. (10 stars, MIT) ## Blogs (1) - [Halfrost-Field](https://madewithwhat.net/vllm/project/halfrost-field/): Source Code Deep Dives, System Design & Engineering Blogs | Halfrost-Field :、 (13,206 stars, CC-BY-SA-4.0) ## Real-time (1) - [xttsv2-vllm-streaming-server](https://madewithwhat.net/vllm/project/xttsv2-vllm-streaming-server/): Real-time streaming TTS server for XTTS-v2 on vLLM — OpenAI-compatible API, ~0.5s TTFB, Docker (54 stars) ## SaaS (1) - [qwen3-asr-mt](https://madewithwhat.net/vllm/project/qwen3-asr-mt/): Multi-tenant streaming ASR server for Qwen3-ASR (vLLM AsyncLLMEngine). Drop-in replacement for qwen-asr-demo-streaming. (13 stars, Apache-2.0)