~ / llm
LLM context
Structured catalog metadata for AI agents and crawlers. Copy the text below or fetch the plain-text endpoint directly.
333 projects · updated 9/5/2026
# Made with vLLM > A curated, daily-updated gallery of the best open-source projects built with vLLM, ranked by GitHub stars. Discover dashboards, UI kits, e-commerce, blogs and dev tools. ## About - Gallery: https://madewithwhat.net/vllm/ - Domain: madewithvllm.com - Projects indexed: 333 - Data source: GitHub (refreshed daily) - Last scraped: 2026-09-06T02:26:46.607099+00:00 ## Top projects - [FunASR](https://madewithwhat.net/vllm/project/funasr/): Open-source speech recognition toolkit for training, inference, streaming ASR, VAD, punctuation, speaker diarization pipelines, and OpenAI-compatible/MCP serving. (19,644 stars, AI & ML) - [llama-cookbook](https://madewithwhat.net/vllm/project/llama-cookbook/): Welcome to the Llama Cookbook! This is your go to guide for Building with Llama: Getting started with Inference, Fine-Tuning, RAG. We also show you how to solve end to end problems using Llama model family and using them on various provider services (18,552 stars, AI & ML) - [Halfrost-Field](https://madewithwhat.net/vllm/project/halfrost-field/): Source Code Deep Dives, System Design & Engineering Blogs | Halfrost-Field :、 (13,214 stars, Blogs) - [AI-Research-SKILLs](https://madewithwhat.net/vllm/project/ai-research-skills/): Comprehensive open-source library of AI research and engineering skills for any AI model. Package the skills and your claude code/codex/gemini agent will be an AI research agent with full horsepower. Maintained by Orchestra Research. (11,405 stars, AI & ML) - [LMCache](https://madewithwhat.net/vllm/project/lmcache/): LMCache: Supercharge Your LLM with the Fastest KV Cache Layer (11,024 stars, AI & ML) - [OpenRLHF](https://madewithwhat.net/vllm/project/openrlhf/): An Easy-to-use, Scalable and High-performance Agentic RL Framework based on Ray (PPO & DAPO & REINFORCE++ & VLM & TIS & vLLM & Ray & Async RL) (9,882 stars, AI & ML) - [inference](https://madewithwhat.net/vllm/project/inference/): Swap GPT for any LLM by changing a single line of code. Xinference lets you run open-source, speech, and multimodal models on cloud, on-prem, or your laptop — all through one unified, production-ready inference API. (9,479 stars, AI & ML) - [dynamo](https://madewithwhat.net/vllm/project/dynamo/): A Datacenter Scale Distributed Inference Serving Framework (7,683 stars, AI & ML) - [Mooncake](https://madewithwhat.net/vllm/project/mooncake/): Mooncake is the serving platform for Kimi, a leading LLM service provided by Moonshot AI. (6,157 stars, AI & ML) - [kserve](https://madewithwhat.net/vllm/project/kserve/): Standardized Distributed Generative and Predictive AI Inference Platform for Scalable, Multi-Framework Deployment on Kubernetes (5,773 stars, AI & ML) - [UltraRAG](https://madewithwhat.net/vllm/project/ultrarag/): A Low-Code MCP Framework for Building Complex and Innovative RAG Pipelines (5,684 stars, AI & ML) - [gpustack](https://madewithwhat.net/vllm/project/gpustack/): A GPU cluster manager for high-performance AI model serving (vLLM, SGLang) and on-demand SSH-accessible GPU instances. (5,436 stars, AI & ML) - [llama-swap](https://madewithwhat.net/vllm/project/llama-swap/): Reliable model swapping for any local OpenAI/Anthropic compatible server - llama.cpp, vllm, etc (5,268 stars, AI & ML) - [sparrow](https://madewithwhat.net/vllm/project/sparrow/): Structured data extraction, instruction calling and agentic workflows with ML, LLM and Vision LLM (5,188 stars, AI & ML) - [semantic-router](https://madewithwhat.net/vllm/project/semantic-router/): A programmable Mixture-of-Models router for heterogeneous LLM inference (5,116 stars, AI & ML) - [tiny-llm](https://madewithwhat.net/vllm/project/tiny-llm/): A course of learning LLM inference serving on Apple Silicon for systems engineers: build a tiny vLLM + Qwen. (4,441 stars, AI & ML) - [LazyLLM](https://madewithwhat.net/vllm/project/lazyllm/): Easiest and laziest way for building multi-agent LLMs applications. (3,860 stars, AI & ML) - [cascadeflow](https://madewithwhat.net/vllm/project/cascadeflow/): Cascading runtime for AI agents. Optimize cost, latency, quality, and policy decisions inside the agent loop. (3,838 stars, AI & ML) - [FastDeploy](https://madewithwhat.net/vllm/project/fastdeploy/): High-performance Inference and Deployment Toolkit for LLMs and VLMs based on PaddlePaddle (3,703 stars, AI & ML) - [Rapid-MLX](https://madewithwhat.net/vllm/project/rapid-mlx/): The fastest local AI engine for Apple Silicon. 4.2x faster than Ollama, 0.08s cached TTFT, 100% tool calling. 17 tool parsers, prompt cache, reasoning separation, cloud routing. Drop-in OpenAI replacement. Works with Claude Code, Cursor, Aider. (3,400 stars, AI & ML) - [ramalama](https://madewithwhat.net/vllm/project/ramalama/): RamaLama is an open-source developer tool that simplifies the local serving of AI models from any source and facilitates their use for inference in production, all through the familiar language of containers. (2,987 stars, AI & ML) - [vllm-ascend](https://madewithwhat.net/vllm/project/vllm-ascend/): Community maintained hardware plugin for vLLM on Ascend (2,560 stars, AI & ML) - [local-studio](https://madewithwhat.net/vllm/project/local-studio/): Control panel for VLLM, Sglang, llama.cpp, exllamav3 (1,557 stars, AI & ML) - [auto-round](https://madewithwhat.net/vllm/project/auto-round/): A SOTA quantization algorithm for high-accuracy low-bit LLM inference, seamlessly optimized for CPU/XPU/CUDA, with multi-datatype support and full compatibility with vLLM, SGLang, and Transformers. (1,552 stars, AI & ML) - [vllm-mlx](https://madewithwhat.net/vllm/project/vllm-mlx/): OpenAI and Anthropic compatible server for Apple Silicon. Run LLMs and vision-language models (Llama, Qwen-VL, LLaVA) with continuous batching, MCP tool calling, and multimodal support. Native MLX backend, 400+ tok/s. Works with Claude Code. (1,488 stars, AI & ML) - [InferenceX](https://madewithwhat.net/vllm/project/inferencex/): Open Source Continuous Inference Benchmark Research Platform — Kimi K3 2.8T, MiniMax M3, DeepSeekv4, GLM5 - GB200 NVL72 vs MI355X vs B200 vs GB300 NVL72 & soon™ TPUv6e/v7/Trainium2/3 | — Kimi K2.7-Code、MiniMax M3、DeepSeekv4、GLM5 - GB200 NVL72 vs MI355X vs B200 vs GB300 NVL72,™ TPUv6e/v7/Trainium2/3 (1,318 stars, AI & ML) - [kubeai](https://madewithwhat.net/vllm/project/kubeai/): AI Inference Operator for Kubernetes. The easiest way to serve ML models in production. Supports VLMs, LLMs, embeddings, and speech-to-text. (1,238 stars, AI & ML) - [GPTQModel](https://madewithwhat.net/vllm/project/gptqmodel/): LLM model quantization (compression) toolkit with HW acceleration support for Nvidia, AMD, Intel GPU and Intel/AMD/Apple CPU via HF, vLLM, and SGLang. (1,222 stars, AI & ML) - [BricksLLM](https://madewithwhat.net/vllm/project/bricksllm/): Enterprise-grade API gateway that helps you monitor and impose cost or rate limits per API key. Get fine-grained access control and monitoring per user, application, or environment. Supports OpenAI, Azure OpenAI, Anthropic, vLLM, and open-source LLMs. (1,221 stars, AI & ML) - [kvcached](https://madewithwhat.net/vllm/project/kvcached/): Virtualized Elastic KV Cache for Dynamic GPU Sharing and Beyond (1,128 stars, AI & ML) ## Useful links - Full catalog, all 333 projects by category: https://madewithwhat.net/vllm/llms-full.txt - RSS feed: https://madewithwhat.net/vllm/rss.xml - Submit a project: https://madewithwhat.net/vllm/submit/ - Browse categories: https://madewithwhat.net/vllm/categories/ - Newsletter: https://madewithwhat.net/vllm/newsletter/ ## Optional This catalog ranks open-source projects by GitHub stars. Each project page includes description, stack, languages, license, and related projects.