~ / llm
LLM context
Structured catalog metadata for AI agents and crawlers. Copy the text below or fetch the plain-text endpoint directly.
llms.txt ↓llms-full.txt ↓RSS feed
289 projects · updated 7/22/2026
# Made with vLLM

> A curated, daily-updated gallery of the best open-source projects built with vLLM, ranked by GitHub stars. Discover dashboards, UI kits, e-commerce, blogs and dev tools.

## About

- Gallery: https://madewithwhat.net/vllm/
- Domain: madewithvllm.com
- Projects indexed: 289
- Data source: GitHub (refreshed daily)
- Last scraped: 2026-07-22T09:29:08.088214+00:00

## Top projects

- [FunASR](https://madewithwhat.net/vllm/project/funasr/): Industrial-grade speech recognition toolkit: 170x realtime, 50+ languages, speaker diarization, emotion detection, streaming, and OpenAI-compatible API. (19,219 stars, AI & ML)
- [llama-cookbook](https://madewithwhat.net/vllm/project/llama-cookbook/): Welcome to the Llama Cookbook! This is your go to guide for Building with Llama: Getting started with Inference, Fine-Tuning, RAG. We also show you how to solve end to end problems using Llama model family and using them on various provider services (18,401 stars, AI & ML)
- [Halfrost-Field](https://madewithwhat.net/vllm/project/halfrost-field/): Source Code Deep Dives, System Design & Engineering Blogs | Halfrost-Field :、 (13,206 stars, Blogs)
- [AI-Research-SKILLs](https://madewithwhat.net/vllm/project/ai-research-skills/): Comprehensive open-source library of AI research and engineering skills for any AI model. Package the skills and your claude code/codex/gemini agent will be an AI research agent with full horsepower. Maintained by Orchestra Research. (10,689 stars, AI & ML)
- [LMCache](https://madewithwhat.net/vllm/project/lmcache/): LMCache: Supercharge Your LLM with the Fastest KV Cache Layer (10,541 stars, AI & ML)
- [OpenRLHF](https://madewithwhat.net/vllm/project/openrlhf/): An Easy-to-use, Scalable and High-performance Agentic RL Framework based on Ray (PPO & DAPO & REINFORCE++ & VLM & TIS & vLLM & Ray & Async RL) (9,785 stars, AI & ML)
- [inference](https://madewithwhat.net/vllm/project/inference/): Swap GPT for any LLM by changing a single line of code. Xinference lets you run open-source, speech, and multimodal models on cloud, on-prem, or your laptop — all through one unified, production-ready inference API. (9,429 stars, AI & ML)
- [dynamo](https://madewithwhat.net/vllm/project/dynamo/): A Datacenter Scale Distributed Inference Serving Framework (7,479 stars, AI & ML)
- [Mooncake](https://madewithwhat.net/vllm/project/mooncake/): Mooncake is the serving platform for Kimi, a leading LLM service provided by Moonshot AI. (5,816 stars, AI & ML)
- [kserve](https://madewithwhat.net/vllm/project/kserve/): Standardized Distributed Generative and Predictive AI Inference Platform for Scalable, Multi-Framework Deployment on Kubernetes (5,681 stars, AI & ML)
- [UltraRAG](https://madewithwhat.net/vllm/project/ultrarag/): A Low-Code MCP Framework for Building Complex and Innovative RAG Pipelines (5,643 stars, AI & ML)
- [gpustack](https://madewithwhat.net/vllm/project/gpustack/): A GPU cluster manager for high-performance AI model serving (vLLM, SGLang) and on-demand SSH-accessible GPU instances. (5,317 stars, AI & ML)
- [sparrow](https://madewithwhat.net/vllm/project/sparrow/): Structured data extraction, instruction calling and agentic workflows with ML, LLM and Vision LLM (5,181 stars, AI & ML)
- [llama-swap](https://madewithwhat.net/vllm/project/llama-swap/): Reliable model swapping for any local OpenAI/Anthropic compatible server - llama.cpp, vllm, etc (4,992 stars, AI & ML)
- [semantic-router](https://madewithwhat.net/vllm/project/semantic-router/): Intelligent Mixture-of-Models Router for Efficient Heterogeneous LLMs Inference (4,962 stars, AI & ML)
- [tiny-llm](https://madewithwhat.net/vllm/project/tiny-llm/): A course of learning LLM inference serving on Apple Silicon for systems engineers: build a tiny vLLM + Qwen. (4,361 stars, AI & ML)
- [LazyLLM](https://madewithwhat.net/vllm/project/lazyllm/): Easiest and laziest way for building multi-agent LLMs applications. (3,853 stars, AI & ML)
- [FastDeploy](https://madewithwhat.net/vllm/project/fastdeploy/): High-performance Inference and Deployment Toolkit for LLMs and VLMs based on PaddlePaddle (3,702 stars, AI & ML)
- [cascadeflow](https://madewithwhat.net/vllm/project/cascadeflow/): Cascading runtime for AI agents. Optimize cost, latency, quality, and policy decisions inside the agent loop. (3,295 stars, AI & ML)
- [Rapid-MLX](https://madewithwhat.net/vllm/project/rapid-mlx/): The fastest local AI engine for Apple Silicon. 4.2x faster than Ollama, 0.08s cached TTFT, 100% tool calling. 17 tool parsers, prompt cache, reasoning separation, cloud routing. Drop-in OpenAI replacement. Works with Claude Code, Cursor, Aider. (3,262 stars, AI & ML)
- [ramalama](https://madewithwhat.net/vllm/project/ramalama/): RamaLama is an open-source developer tool that simplifies the local serving of AI models from any source and facilitates their use for inference in production, all through the familiar language of containers. (2,956 stars, AI & ML)
- [vllm-ascend](https://madewithwhat.net/vllm/project/vllm-ascend/): Community maintained hardware plugin for vLLM on Ascend (2,496 stars, AI & ML)
- [auto-round](https://madewithwhat.net/vllm/project/auto-round/): A SOTA quantization algorithm for high-accuracy low-bit LLM inference, seamlessly optimized for CPU/XPU/CUDA, with multi-datatype support and full compatibility with vLLM, SGLang, and Transformers. (1,519 stars, AI & ML)
- [vllm-mlx](https://madewithwhat.net/vllm/project/vllm-mlx/): OpenAI and Anthropic compatible server for Apple Silicon. Run LLMs and vision-language models (Llama, Qwen-VL, LLaVA) with continuous batching, MCP tool calling, and multimodal support. Native MLX backend, 400+ tok/s. Works with Claude Code. (1,428 stars, AI & ML)
- [local-studio](https://madewithwhat.net/vllm/project/local-studio/): Control panel for VLLM, Sglang, llama.cpp, exllamav3 (1,359 stars, AI & ML)
- [InferenceX](https://madewithwhat.net/vllm/project/inferencex/): Open Source Continuous Inference Benchmark Research Platform — Kimi K2.7-Code, MiniMax M3, DeepSeekv4, GLM5 - GB200 NVL72 vs MI355X vs B200 vs GB300 NVL72 & soon™ TPUv6e/v7/Trainium2/3 | — Kimi K2.7-Code、MiniMax M3、DeepSeekv4、GLM5 - GB200 NVL72 vs MI355X vs B200 vs GB300 NVL72,™ TPUv6e/v7/Trainium2/3 (1,244 stars, AI & ML)
- [kubeai](https://madewithwhat.net/vllm/project/kubeai/): AI Inference Operator for Kubernetes. The easiest way to serve ML models in production. Supports VLMs, LLMs, embeddings, and speech-to-text. (1,222 stars, AI & ML)
- [BricksLLM](https://madewithwhat.net/vllm/project/bricksllm/): Enterprise-grade API gateway that helps you monitor and impose cost or rate limits per API key. Get fine-grained access control and monitoring per user, application, or environment. Supports OpenAI, Azure OpenAI, Anthropic, vLLM, and open-source LLMs. (1,217 stars, AI & ML)
- [GPTQModel](https://madewithwhat.net/vllm/project/gptqmodel/): LLM model quantization (compression) toolkit with HW acceleration support for Nvidia, AMD, Intel GPU and Intel/AMD/Apple CPU via HF, vLLM, and SGLang. (1,205 stars, AI & ML)
- [prometheus-eval](https://madewithwhat.net/vllm/project/prometheus-eval/): Evaluate your LLM's response with Prometheus and GPT4 (1,102 stars, AI & ML)

## Useful links

- Full catalog, all 289 projects by category: https://madewithwhat.net/vllm/llms-full.txt
- RSS feed: https://madewithwhat.net/vllm/rss.xml
- Submit a project: https://madewithwhat.net/vllm/submit/
- Browse categories: https://madewithwhat.net/vllm/categories/
- Newsletter: https://madewithwhat.net/vllm/newsletter/

## Optional

This catalog ranks open-source projects by GitHub stars. Each project page includes description, stack, languages, license, and related projects.
$ made-with-vllm
rssllmmadewithwhat

Ask MadeWithWhat

AI answers may contain mistakes — please double-check important details.