// 289 open-source projects indexed

$ Discover the best projects built with vLLM

A hand-curated, daily-updated gallery of open-source projects. Browse the ecosystem, ranked by GitHub stars.
>
sort --stars
#all#aiml244#devtools35#docs4#dashboards3#blogs1
// 289 results
AI & ML
Industrial-grade speech recognition toolkit: 170x realtime, 50+ languages, speaker diarization, emotion detection, streaming, and OpenAI-compatible API.
@modelscope6d ago
AI & ML
Welcome to the Llama Cookbook! This is your go to guide for Building with Llama: Getting started with Inference, Fine-Tuning, RAG. We also show you how to solve end to end problems using Llama model family and using them on various provider services
@meta-llama6d ago
Blogs
Source Code Deep Dives, System Design & Engineering Blogs | Halfrost-Field :、
@halfrost6d ago
AI & ML
Comprehensive open-source library of AI research and engineering skills for any AI model. Package the skills and your claude code/codex/gemini agent will be an AI research agent with full horsepower. Maintained by Orchestra Research.
@Orchestra-Research6d ago
AI & ML
LMCache: Supercharge Your LLM with the Fastest KV Cache Layer
@LMCache6d ago
AI & ML
An Easy-to-use, Scalable and High-performance Agentic RL Framework based on Ray (PPO & DAPO & REINFORCE++ & VLM & TIS & vLLM & Ray & Async RL)
@OpenRLHF6d ago
AI & ML
Swap GPT for any LLM by changing a single line of code. Xinference lets you run open-source, speech, and multimodal models on cloud, on-prem, or your laptop — all through one unified, production-ready inference API.
@xorbitsai6d ago
AI & ML
A Datacenter Scale Distributed Inference Serving Framework
@ai-dynamo6d ago
AI & ML
Mooncake is the serving platform for Kimi, a leading LLM service provided by Moonshot AI.
@kvcache-ai6d ago
AI & ML
Standardized Distributed Generative and Predictive AI Inference Platform for Scalable, Multi-Framework Deployment on Kubernetes
@kserve6d ago
AI & ML
A Low-Code MCP Framework for Building Complex and Innovative RAG Pipelines
@OpenBMB6d ago
AI & ML
A GPU cluster manager for high-performance AI model serving (vLLM, SGLang) and on-demand SSH-accessible GPU instances.
@gpustack6d ago
AI & ML
Structured data extraction, instruction calling and agentic workflows with ML, LLM and Vision LLM
@katanaml6d ago
AI & ML
Reliable model swapping for any local OpenAI/Anthropic compatible server - llama.cpp, vllm, etc
@mostlygeek6d ago
AI & ML
Intelligent Mixture-of-Models Router for Efficient Heterogeneous LLMs Inference
@vllm-project6d ago
AI & ML
A course of learning LLM inference serving on Apple Silicon for systems engineers: build a tiny vLLM + Qwen.
@skyzh6d ago
AI & ML
Easiest and laziest way for building multi-agent LLMs applications.
@LazyAGI6d ago
AI & ML
High-performance Inference and Deployment Toolkit for LLMs and VLMs based on PaddlePaddle
@PaddlePaddle6d ago
AI & ML
Cascading runtime for AI agents. Optimize cost, latency, quality, and policy decisions inside the agent loop.
@lemony-ai6d ago
AI & ML
The fastest local AI engine for Apple Silicon. 4.2x faster than Ollama, 0.08s cached TTFT, 100% tool calling. 17 tool parsers, prompt cache, reasoning separation, cloud routing. Drop-in OpenAI replacement. Works with Claude Code, Cursor, Aider.
@raullenchai6d ago
AI & ML
RamaLama is an open-source developer tool that simplifies the local serving of AI models from any source and facilitates their use for inference in production, all through the familiar language of containers.
@containers6d ago
AI & ML
Community maintained hardware plugin for vLLM on Ascend
@vllm-project6d ago
AI & ML
A SOTA quantization algorithm for high-accuracy low-bit LLM inference, seamlessly optimized for CPU/XPU/CUDA, with multi-datatype support and full compatibility with vLLM, SGLang, and Transformers.
@intel6d ago
AI & ML
OpenAI and Anthropic compatible server for Apple Silicon. Run LLMs and vision-language models (Llama, Qwen-VL, LLaVA) with continuous batching, MCP tool calling, and multimodal support. Native MLX backend, 400+ tok/s. Works with Claude Code.
@waybarrios6d ago
AI & ML
Control panel for VLLM, Sglang, llama.cpp, exllamav3
@sybil-solutions6d ago
AI & ML
Open Source Continuous Inference Benchmark Research Platform — Kimi K2.7-Code, MiniMax M3, DeepSeekv4, GLM5 - GB200 NVL72 vs MI355X vs B200 vs GB300 NVL72 & soon™ TPUv6e/v7/Trainium2/3 | — Kimi K2.7-Code、MiniMax M3、DeepSeekv4、GLM5 - GB200 NVL72 vs MI355X vs B200 vs GB300 NVL72,™ TPUv6e/v7/Trainium2/3
@SemiAnalysisAI6d ago
AI & ML
AI Inference Operator for Kubernetes. The easiest way to serve ML models in production. Supports VLMs, LLMs, embeddings, and speech-to-text.
@kubeai-project6d ago
AI & ML
Enterprise-grade API gateway that helps you monitor and impose cost or rate limits per API key. Get fine-grained access control and monitoring per user, application, or environment. Supports OpenAI, Azure OpenAI, Anthropic, vLLM, and open-source LLMs.
@bricks-cloud6d ago
AI & ML
LLM model quantization (compression) toolkit with HW acceleration support for Nvidia, AMD, Intel GPU and Intel/AMD/Apple CPU via HF, vLLM, and SGLang.
@ModelCloud6d ago
AI & ML
Evaluate your LLM's response with Prometheus and GPT4
@prometheus-eval6d ago
AI & ML
Virtualized Elastic KV Cache for Dynamic GPU Sharing and Beyond
@ovg-project6d ago
AI & ML
Build your own high performance LLM inference engine in C++ and CUDA - a smaller version of vLLM
@jmaczan6d ago
AI & ML
LLM notes, including model inference, transformer model structure, and llm framework code analysis notes.
@harleyszhang6d ago
AI & ML
Make Discord your LLM frontend - Supports any OpenAI compatible API (OpenRouter, Ollama and more)
@jakobdylanc6d ago
DevTools
UniRL is a Framework for Unified Multimodal Model Reinforcement Learning
@Tencent-Hunyuan6d ago
AI & ML
A Python library powered by Language Models (LLMs) for conversational data discovery and analysis.
@pgalko6d ago
AI & ML
An local, offline (after initial setup), portable OCR software that can process images and PDF files, using DeepSeek-OCR-2 AI (running directly on your machine).
@th1nhhdk6d ago
DevTools
[EMNLP 2024 & AAAI 2026] A powerful toolkit for compressing large models including LLMs, VLMs, and video generative models.
@ModelTC6d ago
AI & ML
Accurate, large-scale, and extensible simulator for LLM inference Systems
@microsoft6d ago
AI & ML
EfficientSAM3 compresses SAM3 into lightweight, edge-friendly models via progressive knowledge distillation for fast promptable concept segmentation and tracking.
@SimonZeng71086d ago
DevTools
SparkTTS、OrpheusTTS,。
@HuiResearch6d ago
AI & ML
AI Bank Statement Document Automation By LLM model and Personal Finanical Analysis
@johnsonhk886d ago
AI & ML
[CVPR 2025] RoboBrain: A Unified Brain Model for Robotic Manipulation from Abstract to Concrete. Official Repository.
@FlagOpen6d ago
Docs
Multimodal RL training framework for diffusion & omni models
@verl-project6d ago
AI & ML
Crater is a cloud-native AI training & inference platform.
@raids-lab6d ago
AI & ML
Pure Rust + CUDA LLM inference engine — no PyTorch, OpenAI-compatible, serves Qwen3 to Kimi-K2
@openinfer-project6d ago
AI & ML
RESTai is an AIaaS (AI as a Service) open-source platform. Supports many public and local LLM suported by Ollama/vLLM/etc. Precise embeddings usage, tuning, analytics etc. Built-in image/audio generation with dynamic loading generators. Live chat deployment. Built-in block based graphical language. Prompt versioning and much more...
@apocas6d ago
DevTools
A command-line interface tool for serving LLM using vLLM.
@Chen-zexi6d ago
AI & ML
A modern web interface for managing and interacting with vLLM servers (www.github.com/vllm-project/vllm). Supports both GPU and CPU modes, with special optimizations for macOS Apple Silicon and enterprise deployment on OpenShift/Kubernetes.
@micytao6d ago
AI & ML
Open Model Engine (OME) — Kubernetes operator for LLM serving, GPU scheduling, and model lifecycle management. Works with SGLang, vLLM, TensorRT-LLM, and Triton
@ome-projects6d ago
AI & ML
The Runpod worker template for serving our large language model endpoints. Powered by vLLM.
@runpod-workers6d ago
AI & ML
KVarN is a native vLLM KV-cache quantization backend for your agents: 3-5x more context, throughput above FP16, and FP16-level accuracy. Calibration-free, one flag.
@huawei-csl6d ago
AI & ML
Ready-to-use DeepSeek-OCR Web UI | Modern Interface | 7 Recognition Modes | Batch Processing | Real-time Logging | Fully Responsive
@neosun1006d ago
AI & ML
Self-hosted AI agent OS. Your memory, chat, agents, and files stay on hardware you own, offline by default, cloud by choice. Offline AI memory (taOSmd), self-hosted multi-framework group chat, a full web desktop + app store, and auto-clustering across the consumer hardware you already have (Orange/Raspberry Pi, Mac mini, gaming PC).
@jaylfc6d ago
AI & ML
The open creative AI workspace
@nodetool-ai6d ago
AI & ML
Fully uncensored, capability-enhanced abliteration of Qwen3.6-27B. NVFP4 + z-lab DFlash speculative decoding (n=12) on the unified ghcr.io/aeon-7/aeon-vllm-ultimate:latest container, tuned for long-context draft acceptance on DGX Spark. 6 HF variants (BF16/NVFP4/MTP/MTP-XS), docker-compose, and QuickStart.
@AEON-76d ago
AI & ML
TopicGPT: A Prompt-Based Framework for Topic Modeling [NAACL'24]
@chtmp2236d ago
AI & ML
Python package for LLM compression
@FujitsuResearch6d ago
AI & ML
Low latency JSON generation using LLMs
@varunshenoy6d ago
AI & ML
Engine-agnostic LLM gateway in Rust. Full OpenAI & Anthropic API compatibility across vLLM, TRT-LLM, TokenSpeed, SGLang, OpenAI, Gemini & more. Industry-first gRPC pipeline, KV cache-aware routing, chat history, tokenization caching, Responses API, embeddings, WASM plugins, MCP, and multi-tenant auth.
@lightseekorg6d ago
AI & ML
LvLLM is a special NUMA extension of vllm that makes full use of CPU and memory resources, reduces GPU memory requirements, and features an efficient GPU parallel and NUMA parallel architecture, supporting hybrid inference for MOE large models.
@guqiong966d ago
AI & ML
sparkrun - launch, manage, and stop LLM inference workloads on NVIDIA DGX Spark systems
@spark-arena6d ago
AI & ML
Setup and run a local LLM and Chatbot using consumer grade hardware.
@jasonacox6d ago
AI & ML
Easy, advanced inference platform for large language models on Kubernetes. Star to support our work!
@InftyAI6d ago
Docs
Implementation for FP8/INT8 Rollout for RL training without performence drop.
@yaof206d ago
AI & ML
Persist and reuse KV Cache to speedup your LLM.
@ModelEngine-Group6d ago
DevTools
AI-powered offensive security testing using autonomous agents, directly in your terminal.
@pensarai6d ago
AI & ML
The production framework for Predictive and Generative AI. Serve any model as an API in one line, with OpenAI/Anthropic/Ollama-compatible endpoints, a built-in chat UI, and native MCP.
@vortico6d ago
AI & ML
Blazing-fast LLM inference in pure Rust. No PyTorch and Python runtime.
@guoqingbao6d ago
DevTools
ML software (llama.cpp, ComfyUI, vLLM) builds for AMD gfx906 GPUs, e.g. Radeon VII / MI50 / MI60
@mixa36076d ago
DevTools
Qwen3.5-122B-A10B on DGX Spark: 28.3 → 51 tok/s (+80%)
@albond6d ago
AI & ML
LLM ,。
@lework6d ago
AI & ML
[ICML-2026] Official implementation of "SEAgent: Self-Evolving Computer Use Agent with Autonomous Learning from Experience"
@SunzeY6d ago
AI & ML
High-performance lightweight proxy and load balancer for LLM infrastructure. Intelligent routing, automatic failover and unified model discovery across local and remote inference backends.
@thushan6d ago
AI & ML
A CPU Realtime VLM in 500M. Surpassed Moondream2 and SmolVLM. Training from scratch with ease.
@lucasjinreal6d ago
AI & ML
An Open-source, self-hosted AI model hub with Hugging Face compatibility, accelerating vLLM/SGLang performance.
@matrixhub-ai6d ago
AI & ML
gpt_serverLLMs、Embedding、Reranker、ASR、TTS、、。
@shell-nlp6d ago
AI & ML
Home Assistant LLM integration for local OpenAI-compatible services (llamacpp, vllm, etc)
@skye-harris6d ago
AI & ML
One-click Qwen3.6-27B inference on Windows. 158 tok/s on RTX 5090, 72 tok/s on RTX 3090. Native, no WSL, no Docker, no telemetry.
@devnen6d ago
AI & ML
a fun and educational take on vLLM
@ovshake6d ago
AI & ML
A PyTorch native library for training speculative decoding models
@lightseekorg6d ago
AI & ML
A general-purpose API load testing platform that supports LLM services and business HTTP interfaces, enabling one-click performance testing, result comparison, and AI-powered intelligent analysis and summarization. API , HTTP ,、 AI
@MigoXLab6d ago
AI & ML
Open-source tools for training and evaluating Vision Language Models for OCR
@Roots-Automation6d ago
AI & ML
Fully-featured, beautiful web interface for vLLM - built with NextJS.
@yoziru6d ago
Docs
vLLM Documentation in Chinese Simplified / vLLM
@hyperai6d ago
AI & ML
[ICCV2025] Referring any person or objects given a natural language description. Code base for RexSeek and HumanRef Benchmark
@IDEA-Research6d ago
AI & ML
High-performance KV cache storage for LLM inference — GPU offloading, SSD caching, and cross-node sharing via RDMA. Works with vLLM and SGLang.
@novitalabs6d ago
AI & ML
Booster - open accelerator for LLM models. Better inference and debugging for AI hackers
@gotzmann6d ago
AI & ML
Deep Learning Deployment Framework: Supports tf/torch/trt/trtllm/vllm and other NN frameworks. Support dynamic batching, and streaming modes. It is dual-language compatible with Python and C++, offering scalability, extensibility, and high performance. It helps users quickly deploy models and provide services through HTTP/RPC interfaces.
@NetEase-Media6d ago
AI & ML
Probing Scientific General Intelligence of LLMs with Scientist-Aligned Workflows
@InternScience6d ago
AI & ML
Kubernetes operator for self-hosted LLM inference across a heterogeneous GPU fleet: NVIDIA CUDA, AMD Vulkan, and Apple Silicon Metal. Runtimes: llama.cpp, vLLM, TGI, mlx-server. Multi-GPU sharding, model caching, OpenAI-compatible endpoints. Apache-2.0, run across homelab and on-prem fleets, actively developed.
@defilantech6d ago
AI & ML
CacheRoute is an innovative LLM scheduling scheme dedicated to enabling flexible KV cache reuse across LLM systems, improving task performance and system efficiency.
@BJTU-ANT6d ago
AI & ML
Automated system for LLM evaluation via agents. Doc as below:
@OpenDCAI6d ago
AI & ML
Cognithor · Agent OS: Local-first autonomous agent operating system. 19 LLM providers, 18 channels, 145 MCP tools, 6-tier memory, Agent Packs marketplace, zero telemetry. Python 3.12+, Apache 2.0.
@Alex8791-cyber6d ago
AI & ML
GLM-5.2-NVFP4-REAP-469B serving on SM120 (4× RTX PRO 6000 Blackwell) — one-command vLLM launch recipe, 250K context, DeepSeek Sparse Attention + MTP speculative decode
@0xSero6d ago
DevTools
vLLM-5090: Docker Container for RTX 5090 + OpenCode
@BoltzmannEntropy6d ago
AI & ML
DocMind AI is a powerful, open-source Streamlit application leveraging LlamaIndex, LangGraph, and local Large Language Models (LLMs) via Ollama, LMStudio, llama.cpp, or vLLM for advanced document analysis. Analyze, summarize, and extract insights from a wide array of file formats, securely and privately, all offline.
@BjornMelin6d ago
AI & ML
Unified management and routing for llama.cpp, MLX and vLLM models with web dashboard.
@lordmathis6d ago
AI & ML
Never give AI companies your secrets! A local LLM-based privacy filter for LLM users. Seamless integration with your existing AI tools as a Python library / OpenAI SDK replacement / API Gatetway / Web Server.
@cxumol6d ago
AI & ML
A project worth exploring.
@GaeaRuiW6d ago
AI & ML
SNDR Core Engine (Genesis) — vLLM runtime patch-overlay for Qwen3.6 + Gemma4 on consumer NVIDIA (Ampere sm_86, 2× A5000/3090). Qwen3.6-35B-A3B FP8 ~240 tok/s, 27B-int4 hybrid GDN+Mamba, Gemma4 26B/31B AWQ, 256K ctx. 321 patches: TurboQuant k8v4 KV, MTP/DFlash spec-decode, FULL cudagraph, hybrid GDN. vLLM pin dev424 + Control Center GUI.
@Sandermage6d ago
AI & ML
Privacy-first LLM proxy and AI gateway — load balancing, multi-provider routing, API key management, usage tracking, rate limiting. Self-hosted. Zero knowledge of your prompts.
@voidmind-io6d ago
AI & ML
A High-Performance LLM Inference Engine with vLLM-Style Continuous Batching
@lumia4316d ago
AI & ML
Convert PowerPoint files into semantically rich text using vision language models
@ALucek6d ago
DevTools
KEDA External gRPC Scaler for GPU workloads - native NVML metrics via DaemonSet, no Prometheus required
@pmady6d ago
AI & ML
Efficient LLM inference on Slurm clusters.
@VectorInstitute6d ago
AI & ML
One-command vLLM installation for NVIDIA DGX Spark with Blackwell GB10 GPUs (sm_121 architecture)
@eelbaz6d ago
AI & ML
llm-inference is a platform for publishing and managing llm inference, providing a wide range of out-of-the-box features for model deployment, such as UI, RESTful API, auto-scaling, computing resource management, monitoring, and more.
@OpenCSGs6d ago
AI & ML
Kubernetes-native platform for deploying and managing AI inference across multiple providers
@kaito-project6d ago
AI & ML
Extensible generative AI platform on Kubernetes with OpenAI-compatible APIs.
@llmariner6d ago
AI & ML
FlowSteer: agents designing agentic workflows via reinforced progressive canvas editing.
@beita69696d ago
AI & ML
Nano vLLM with vLLM v1's request scheduling strategy and chunked prefill
@slwang-ustc6d ago
AI & ML
Qwen2+SFT+DPO, SFTTrainer/DPOTrainer/TRPOTrainer,,(neo4j, milvus, LDA, )。, vllm, vllm API embedder + Reranker RAG 。 MDAgents , vllm api 。
@NJUxlj6d ago
DevTools
Local diagnostic CLI for NVIDIA DGX Spark (GB10). Detects power caps, UMA pressure, thermal risk, CUDA 13/SM_121 wheel mismatches, Docker issues, and validates vLLM/Ollama/llama.cpp/SGLang recipes.
@joeynyc6d ago
AI & ML
Enterprise-grade LLM automated deployment tool that makes AI servers truly "plug-and-play".
@EM-GeekLab6d ago
AI & ML
This repository contains hands-on projects, code examples, and deployment workflows. Explore multi-agent systems, LangChain, LangGraph, AutoGen, CrewAI, RAG, MCP, automation with n8n, and scalable agent deployment using Docker, AWS, and BentoML.
@MDalamin56d ago
AI & ML
A unified evaluation toolkit and leaderboard for rigorously assessing the scientific intelligence of large language and vision–language models across the full research workflow.
@InternScience6d ago
AI & ML
Granite Switch — Build AI models like you build software
@generative-computing6d ago
AI & ML
Practical local LLM recipes and benchmarks for RTX 5060 Ti setups
@5p00kyy6d ago
AI & ML
A simple service that integrates vLLM with Ray Serve for fast and scalable LLM serving.
@asprenger6d ago
AI & ML
A High-Efficiency System of Large Language Model Based Search Agents
@tiannuo-yang6d ago
AI & ML
A comprehensive toolkit for deploying production-ready Generative AI infrastructure on Amazon EKS. Includes pre-configured components for: AI Gateway (LiteLLM) LLM Serving (vLLM, SGLang, Ollama) Vector Databases, Embedding Models (TEI) Observability (Langfuse, Phoenix) etc. Fast-track your GenAI deployment with Kubernetes
@aws-samples6d ago
AI & ML
First open-source implementation of Google TurboQuant (ICLR 2026) -- near-optimal KV cache compression for LLM inference. 5x compression with near-zero quality loss.
@OnlyTerp6d ago
AI & ML
Simple LLM service identification - translate IP:Port to Ollama, vLLM, LiteLLM, or 60+ other AI services in seconds
@praetorian-inc6d ago
AI & ML
Zero trust LLM gateway. OpenAI-compatible proxy with semantic routing and load balancing across OpenAI, Anthropic, Ollama, vLLM, and any compatible backend. Identity-based access, virtual API keys, and end-to-end encryption via OpenZiti
@openziti6d ago
AI & ML
Soup turns the pain of LLM fine-tuning into a simple workflow. One config, one command, done.
@MakazhanAlpamys6d ago
AI & ML
Reliable and Efficient Semantic Prompt Caching with vCache
@vcache-project6d ago
Dashboards
Real-time hardware and LLM inference monitoring — GPU, CPU, memory, and vLLM metrics streamed to a dashboard.
@niklasfrick6d ago
AI & ML
Production inference for encoder models - ColBERT, GLiNER, ColPali, embeddings etc. - as vLLM plugins for online and in-process deployment
@latenceainew6d ago
AI & ML
[ICML 2026] Decoding Tree Sketching (DTS): a training-free & model agonistic & plug-in framework for LLM parallel reasoning.
@ZichengXu6d ago
AI & ML
Agent and LLM API performance monitoring via network packet probe. Measures performance of OpenClaw, Claude, Codex, DeepAgents and more — deployed on the provider side, no SDK changes required.
@Netis6d ago
AI & ML
PDF Parsing Tool: GOT's vLLM acceleration implementation, MinerU for layout recognition, and GOT for table formula parsing.
@liunian-Jay6d ago
AI & ML
Automated Deep Research with LLMs, web search, paper parsing, and didactic summarization.
@protonspy6d ago
AI & ML
This project aims to provide a high effective KV cache manage framework for llm inference and improve memory utilization and inference speed.
@TheToughCrane6d ago
Dashboards
An intelligent tuner for vLLM that automatically monitors GPU metrics, uses Bayesian optimization to tune parameters
@jranaraki6d ago
AI & ML
EmbodiedAgents is a fully-loaded ROS2 based framework for creating interactive physical agents that can understand, remember, and act upon contextual information from their environment.
@automatika-robotics6d ago
AI & ML
LLM-Inference-Bench
@argonne-lcf6d ago
DevTools
Sardeenz is a proof-of-concept application that allows you to load more than one model on a given GPU. It allows you to add more and more models onto a GPU, until it is fully utilized.
@rh-aiservices-bu6d ago
AI & ML
An endpoint server for efficiently serving quantized open-source LLMs for code.
@wangcx186d ago
AI & ML
Modern AI chatbot supporting multiple LLMs. Switch between Gemini, Mistral, Llama, Claude and ChatGPT.
@intelligentnode6d ago
AI & ML
LLM-powered security log analyzer: detect threats & anomalies with zero regex — just declare a Pydantic schema. Real-time Telegram alerts, SIEM-ready with Elasticsearch/Kibana. Supports OpenAI, Ollama, vLLM.
@call5186d ago
AI & ML
A tool for benchmarking LLMs on Modal
@modal-labs6d ago
AI & ML
vLLM Router
@llm-semantic-router6d ago
AI & ML
High-Performance KV Cache Storage Engine on CXL Shared Memory for LLM Inference
@xcena-dev6d ago
AI & ML
OpenAI-compatible multilingual TTS server — Chatterbox on vLLM with real-time PCM audio streaming, low time-to-first-byte (~0.7 s), voice cloning, and 23 languages.
@wuxuedaifu6d ago
Real-time
Real-time streaming TTS server for XTTS-v2 on vLLM — OpenAI-compatible API, ~0.5s TTFB, Docker
@wuxuedaifu6d ago
AI & ML
Turn documents into structured JSON with local-first document AI. Run 100% locally by default, with API, CLI, and Web UI.
@parsehawk6d ago
AI & ML
DFlash vLLM for DGX Spark — Plug & Play Block-Diffusion Speculative Decoding
@AEON-76d ago
AI & ML
[ICPP'25] TD-Pipe: Temporally-Disaggregated Pipeline Parallelism Architecture for High-Throughput LLM Inference
@MLSysU6d ago
AI & ML
DGX Spark / GB10 vLLM Docker stack for large-model serving, presets, patches, and validation notes.
@bjk1106d ago
AI & ML
Arks is a cloud-native inference framework running on Kubernetes
@scitix6d ago
AI & ML
The open-source AI platform for enterprises that can't send data to the cloud. OpenAI-compatible API, full management dashboard, zero data egress.
@xinity-ai6d ago
AI & ML
Batch Deployment for Document Parsing with AWS Batch & Qwen-2.5-VL
@jeremyarancio6d ago
AI & ML
A high-performance RDMA distributed file system for fast LLM Inference and GPU Training.
@blackbird-io6d ago
AI & ML
TurboQuant KV cache compression plugin for vLLM — asymmetric K/V, 8 models validated, consumer GPUs
@Alberto-Codes6d ago
DevTools
A physics-grounded, cost-aware optimizer for vLLM.
@jungledesh6d ago
AI & ML
vLLM Qwen 3.6-27B (AWQ-INT4) + DFlash speculative decoding on AMD Strix Halo (gfx1151 iGPU, 128 GB UMA, ROCm 7.13). 24.8 t/s single-stream, vision, tool calling, 256K context, OpenAI-compatible, Docker. Matches DGX Spark FP8+DFlash+MTP at a third of the cost. No CUDA.
@hec-ovi6d ago
AI & ML
Native macOS menu bar app for realtime dictation with optional LLM polishing. Connects to any OpenAI Realtime-compatible backend — fully local on Apple Silicon with voxmlx + mlx-swift-lm.
@T0mSIlver6d ago
AI & ML
ElasticMM: Elastic and Efficient MLLM Serving System
@hpdps-group6d ago
DevTools
useful Gentoo overlay Curated ebuilds, AI, tools & science
@istitov6d ago
AI & ML
Local-first CLI for benchmarking LLMs on real hardware — quality, speed, reliability, and a real multi-turn agent loop.
@outsourc-e6d ago
Docs
Synthetic medical VQA pipeline: 119K images annotated by frontier VLMs, cross-validated at 93% agreement, fine-tuned on 3 model families (2-3B params)
@openmed-labs6d ago
AI & ML
Private LLM/RAG platform in one command for NVIDIA DGX Spark / GB10 (arm64). Validated on real hardware.
@botAGI6d ago
AI & ML
Carbon Limiting Auto Tuning for Kubernetes
@Climatik-Project6d ago
AI & ML
Self-hosted, OpenAI-compatible inference for the agentic era: reasoning LLMs, universal tool calling, and the Responses API alongside embeddings, speech, and image models — many models sharing your GPUs, one gateway. Powered by Ray Serve.
@alez0076d ago
DevTools
Code for building self-expanding knowledge graphs with Outlines, vLLM, neo4j, and Modal.
@just-cameron6d ago
AI & ML
Android keyboard with local AI (Ollama, Whisper, MCP) or cloud (Gemini, Groq, OpenAI)
@SvReenen6d ago
AI & ML
Hit your limit? Need privacy? Just swap the model, everything else stays
@luongnv896d ago
AI & ML
AIfred-Intelligence — self-hosted Multi-Agent Assistant with Debate Modes (Symposion/Tribunal), Voice (STT + Streaming-TTS), RAG with Long-Term Memory, Web Research and Tool-Calling. Reachable via Web-UI or Telegram/Discord/Email/EPIM. Multi-Backend: llama-swap, Ollama, vLLM, TabbyAPI.
@Peuqui6d ago
DevTools
Serve Llama 3.3 70B (with AWQ quantization) using vLLM and deploy it on BentoCloud.
@kingabzpro6d ago
DevTools
Benchmarking Open-Ended Inference Optimization by AI Agents
@aisa-group6d ago
AI & ML
Bleeding edge vLLM Docker image for the NVIDIA DGX Spark (GB10 / sm_121a).
@timothystewart66d ago
AI & ML
Lightweight proxy for LLM
@yatesdr6d ago
DevTools
vllm
@gameofdimension6d ago
AI & ML
From-scratch C++/CUDA inference engine for the NVIDIA RTX 5090 (sm_120a) — the best single-GPU backend for agentic AI: tool calling, long-context loops, reasoning and concurrent sub-agents on top of the fastest single-stream decode on the 5090 (beats llama.cpp, at-or-ahead of vLLM on NVFP4). 100% written by Claude Code.
@kekzl6d ago
AI & ML
AI Coding Assistant CLI for offline enterprise environments - Local LLM platform with Plan & Execute architecture, Supervised Mode, and auto-update system
@HanSyngha6d ago
AI & ML
Open-source, self-hostable social listening.
@obris-dev6d ago
DevTools
Policy-driven seamless lazy loading
@cloudpilot-ai6d ago
AI & ML
A REST API for vLLM, production ready
@France-Travail6d ago
AI & ML
a simple lightweight large language model pipeline framework.
@sherlockchou866d ago
AI & ML
Powerful no-code LLM fine-tuner: upload data → train → deploy in minutes. Unsloth 2-5× acceleration · QLoRA/DPO/RLHF/PPO/ORPO · Reward Model training · GGUF export · vLLM inference · BLEU/ROUGE/BERTScore · full CLI · Heretic Mode to unlock full model potential
@Yog-Sotho6d ago
AI & ML
[CVPR2025] Official Repository for IMMUNE: Improving Safety Against Jailbreaks in Multi-modal LLMs via Inference-Time Alignment
@itsvaibhav016d ago
AI & ML
Native Windows vLLM 0.24.0: Python 3.13, CUDA 12.8, PyTorch 2.11 cu128, RTX 30/40/50 wheel, hardened portable installer with legacy PowerShell support, Qwen3-VL/FlashAttention fixes, Rust frontend/tool parser, OpenAI server fixes, and 10 KV-cache compression dtypes.
@aivrar6d ago
AI & ML
Dockerized LLM inference server with constrained output (JSON mode), built on top of vLLM and outlines. Faster, cheaper and without rate limits. Compare the quality and latency to your current LLM API provider.
@phospho-app6d ago
AI & ML
A "standard library" of Triton kernels.
@stackav-oss6d ago
AI & ML
A web-based memory usage and performance calculator for Huggingface GGUF models
@gdevenyi6d ago
AI & ML
Official implementation for Text Generation Beyond Discrete Token Sampling
@EvanZhuang6d ago
AI & ML
Patched native-Windows build of vLLM. Three Windows-specific fixes (CPU-relay for Gloo, Qwen3 reasoning parser, wildcard model name) on top of SystemPanic 0.19.0. Engine for devnen/qwen3.6-windows-server.
@devnen6d ago
AI & ML
A hybrid router that uses Spot GPU instances to reduce costs and Serverless GPUs for making Cold Starts faster.
@Tandemn-Labs6d ago
AI & ML
Orpheus TTS Server with streaming support (TTFB ~160ms)
@taresh186d ago
AI & ML
Python SDK for Agent Vector Protocol – transfer KV-cache between LLM agents instead of text
@VectorArc6d ago
AI & ML
Training-free KV cache compression via E8 lattice VQ. 2-bit KV that preserves retrieval (30/30 NIAH vs TurboQuant 0/30). Calibration-free, 9 architectures validated.
@jagmarques6d ago
AI & ML
Unified KV cache compression for LLM inference — TurboQuant, IsoQuant, PlanarQuant, TriAttention. 10 methods, GPU-validated, multi-GPU planner. Compress KV cache 5-80x to run bigger models, longer context, more agents on your GPU.
@aivrar6d ago
AI & ML
This project is the backend engine for a fully autonomous AI-powered call center. It integrates a large language model (LLM), speech recognition, and text-to-speech to manage real-time phone conversations via Asterisk.
@klikz-dev6d ago
AI & ML
Local AI workstation — discover, run, chat, benchmark, and generate images from open-weight models. DFlash/DDTree speculative decoding, TurboQuant & TriAttention cache compression strategies, MLX + llama.cpp + vLLM + MTPLX backends.
@cryptopoly6d ago
AI & ML
Probing the limitations of multimodal language models for chemistry and materials research
@lamalab-org6d ago
AI & ML
, LLM
@zRzRzRzRzRzRzR6d ago
AI & ML
Run the LLM Swarm Router on machines to distribute Local Ai to the Swarm - More Machines - MORE SPEED
@matthewdcage6d ago
AI & ML
Dual-engine (llama.cpp + vLLM) LLM benchmarking pipeline for GGUF & safetensors on NVIDIA GPUs — speed, quality, live dashboard, publishable cards.
@notwitcheer6d ago
AI & ML
Rust SDK for building AI agents with local OpenAI-compatible servers (LMStudio, Ollama, llama.cpp, vLLM). Features streaming, tools, hooks, retry logic, and comprehensive examples.
@slb3506d ago
AI & ML
Enterprise-grade platform to generate and execute Cypress, Playwright, WebdriverIO, and Appium end-to-end tests from natural language requirements. (Appium is experimental and requires external mobile infrastructure.)
@aiqualitylab6d ago
AI & ML
Say 'dreamier' and your ComfyUI workflow shifts — instantly, reversibly. An AI co-pilot for VFX artists: zero-LLM recipes (dreamier/sharper/faster), 133 MCP tools, EXR-aware vision, workflow.lock provenance, full undo, native sidebar, 5 swappable brains (Claude, GPT, Gemini, Ollama, Nemotron).
@JosephOIbrahim6d ago
DevTools
DeepSeek-V3, R1 671B on 8xH100 Throughput Benchmarks
@dzhsurf6d ago
AI & ML
An 18 notebook course that isolates and measures each component of agentic loop engineering on real, industry standard software datasets.
@FareedKhan-dev6d ago
AI & ML
An imperative command-line-interface for AI workload orchestration
@theoddden6d ago
AI & ML
OpenAI-compatible, vLLM-served OCR API for the Surya-OCR-2 model — multilingual document OCR (layout + text recognition) with request batching, a local CLI, and Docker packaging.
@wuxuedaifu6d ago
AI & ML
[ACL 2026] CoCoA: Collaborative Chain-of-Agents for Parametric-Retrieved Knowledge Synergy
@liunian-Jay6d ago
DevTools
GLM-5.2 (744B/40B MoE) on a 4× DGX Spark / GB10 (sm_121) cluster: portable Triton sparse-MLA kernels, a data-free expert prune, MTP draft, and a one-script bootstrap.
@CosmicRaisins6d ago
AI & ML
Run the AEON Bench suite on your own hardware: verified HuggingFace pull → serve → benchmark (text · agentic ×3 harnesses · vision · audio · arena · perf) → ed25519-signed attested submit.
@AEON-76d ago
AI & ML
NVFP4 inference on Blackwell GeForce (RTX 5090/5080/5070 Ti/RTX PRO 6000) — SM120 patches for vLLM + FlashInfer + CUTLASS. 175 tok/s on Qwen3.6-35B MoE.
@lna-lab6d ago
AI & ML
Self-hosted meeting transcription portal — speech-to-text, speaker diarization, LLM-corrected transcripts, structured summaries and Word minutes, on your own GPUs. Flask + PostgreSQL, GDPR audit trail, distributed GPU topologies, docker
@Martossien6d ago
AI & ML
Why is LLM inference slow — and how do you make it fast? A hands-on, first-principles course: roofline → KV cache → quantization → parallelism → vLLM/SGLang, with GPU labs on open models.
@jiahongsigma6d ago
AI & ML
Low-Cost Cross-Domain Web Structured Information Extraction using specialized LoRA adapters.
@abdo-Mansour6d ago
DevTools
[AAAI 2025 oral] Attribution Analysis Meets Model Editing: Advancing Knowledge Correction in Vision Language Models with VisEdit
@qizhou0006d ago
AI & ML
RAG using LlamaIndex:Computer Network Q&A System powered by LlamaIndex | LlamaIndex - HyDE+ + vLLM +Ragas
@userHanlh6d ago
AI & ML
Comprehensive, scalable ML inference architecture using Amazon EKS, leveraging Graviton processors for cost-effective CPU-based inference and GPU instances for accelerated inference. Guidance provides a complete end-to-end platform for deploying LLMs with agentic AI capabilities, including RAG and MCP
@aws-solutions-library-samples6d ago
AI & ML
ICE-PIXIU:A Cross-Language Financial Megamodeling Framework
@YY06496d ago
AI & ML
A Scheduler for Batched LLM Inference
@iPieter6d ago
AI & ML
Intelligent load balancer for distributed vLLM server clusters vLLM
@xerrors6d ago
AI & ML
A lightweight post-training framework for LLMs and VLMs. 51 algorithms, 38 verified models. Scales with DeepSpeed, vLLM, and Ray.
@warlockee6d ago
AI & ML
Source code for the paper: Hear Both Sides: Efficient Multi-Agent Debate via Diversity-Aware Message Retention
@DA2I2-SLM6d ago
AI & ML
Collection of recent advanced RAG techniques.
@hienhayho6d ago
AI & ML
Object-storage-native KV cache for LLM inference & RL. Cross-restart, cross-conversation, cross-engine via shared S3 bucket.
@Venkat28116d ago
AI & ML
Local-first desktop assistant for the whole job hunt — scrape listings, match with a local LLM, generate applications in your own voice, and track everything on a Kanban board.
@Keljian6d ago
DevTools
DGX Spark research and tests - containers, benchmarks, and investigation notes for running models on GB10 (SM 12.1)
@Sggin16d ago
AI & ML
Drop-in OIDC & Google A2A auth + Weaviate memory for Ollama, vLLM and any local LLM server.
@attach-dev6d ago
AI & ML
Layered prefill changes the scheduling axis from tokens to layers and removes redundant MoE weight reloads while keeping decode stall free. The result is lower TTFT, lower end-to-end latency, and lower energy per token without hurting TBT stability.
@scale-snu6d ago
AI & ML
Manalyzer: End-to-end Automated Meta-analysis with Multi-agent System
@black-yt6d ago
AI & ML
The official repo for the paper "Optimizing LLM Inference Throughput via Memory-aware and SLA-constrained Dynamic Batching"
@KevinLee11106d ago
AI & ML
An OpenAI Compatible API which integrates LLM, Embedding and Reranker. LLM、Embedding Reranker OpenAI API
@hcd2336d ago
AI & ML
KONASH: Train knowledge agents that search, retrieve, and reason. Based on KARL (Databricks, 2026).
@konaequity6d ago
AI & ML
Diagnose vLLM inference servers
@aminalaee6d ago
DevTools
htop for your LLM inference cluster
@InfraWhisperer6d ago
AI & ML
Agentic SDLC with local LLM's.
@xencon6d ago
AI & ML
An open-source, model-agnostic agent harness for local LLMs. Define agents in YAML (tools, memory, deny-first permissions) and run them against any OpenAI-compatible endpoint: vLLM, Ollama, LM Studio, or llama.cpp.
@ahwurm6d ago
AI & ML
Dual-mode cloud burst LLM router. Edge-first or cloud-first inference with data sovereignty, cost controls, and automatic failover. https://aiburstcloud.com
@aiburstcloud6d ago
AI & ML
Serve GLM-5.2 469B (REAP-pruned, NVFP4) across 3× NVIDIA DGX Spark with vLLM pipeline parallelism — 256K context, production-ready config and patches
@bird6d ago
DevTools
The official implementation of NeurlPS 2025 D&B paper: IndustryEQA: Pushing the frontiers of Embodied Question Answering in Industrial Scenarios.
@JackYFL6d ago
AI & ML
Convert pdf and image files into markdown
@arc536d ago
AI & ML
Interested in running a conspiracy of agents (technical term) on your Local AI Infra? Who isn't!
@digitalspaceport6d ago
AI & ML
Optimized vLLM setup for Qwen3.6-27B-FP8 on dual RTX PRO 6000 Blackwell (192 GB GDDR7, no NVLink); config, benchmark sweep results, and custom chat template with thinking mode off by default.
@theogravity6d ago
AI & ML
Easily boost the speed of pulling your models and datasets from various of inference runtimes. (e.g. HuggingFace, Ollama, vLLM, and more!)
@moeru-ai6d ago
AI & ML
vLLM + Qwen3.6-27B (BF16) OpenAI-compatible inference server on AMD Strix Halo (Ryzen AI Max+ 395, gfx1151). Vision input, 256K context, /v1/responses with separated reasoning, via TheRock ROCm.
@hec-ovi6d ago
AI & ML
A hands‑on RAG experimentation lab. Largely configurable with debug insights. Classification‑driven corpus construction, filter chains, document loading, chat interaction, Open WebUI integration. Experimental by design and not production‑ready.
@HarinezumIgel22h ago
DevTools
Reproducible recipe: serve abliterated Gemma-4-12B (gemma4_unified) at 50-118 tok/s on no-NVLink Blackwell (SM120) via vLLM nightly + ModelOpt FP8/NVFP4 + MTP spec-decode.
@lna-lab6d ago
DevTools
A project worth exploring.
@ampir-nn6d ago
AI & ML
chenxuniu/TokenSpark-Benchmark-Benchmarking-Power-Consumption-of-LLM-Inference-on-Multi-Node-Clusters
@chenxuniu6d ago
AI & ML
Batch LLM Inference with Ray Data LLM: From Simple to Advanced
@0-mostafa-rezaee-06d ago
AI & ML
Turn any NVIDIA GPU into a local AI platform. Inference + fine-tuning in your browser. One command to start, automatic clustering.
@getainode6d ago
AI & ML
LLM, Fine Tuning, Llama 2, Gemma, Mixtral, vLLM, LangChain, RAG, ChromaDB, FAISS
@joydeb286d ago
DevTools
Running Large Language Model easily.
@janelu96d ago
AI & ML
FlashHead: Efficient Drop-In Replacement for the Classification Head in Language Model Inference
@embedl6d ago
AI & ML
Exocomp Agentic Environment for Go
@cookiengineer6d ago
AI & ML
Kubernetes scanner that discovers LLMs running on vLLM and extracts their deployment and runtime facts.
@paralleliq6d ago
SaaS
Multi-tenant streaming ASR server for Qwen3-ASR (vLLM AsyncLLMEngine). Drop-in replacement for qwen-asr-demo-streaming.
@jayter-official6d ago
AI & ML
GPU Service Manager for LLM workloads on Linux/NVIDIA systems.
@jaigouk6d ago
AI & ML
SAM — Smart Agentic Model: CLI coding agent for open-source LLMs. pip install sam-agent
@SecFathy6d ago
DevTools
Automated Triton w8a8 block FP8 kernel tuning tool for vLLM. Auto-detects model architecture, supports Qwen3-Coder-30B-A3B-Instruct-FP8/DeepSeek-V3/custom models, multi-GPU parallel tuning, and generates optimized kernel configs for quantization.
@massif-016d ago
AI & ML
A lean, local AI research agent framework that runs on your machine
@KabakaWilliam6d ago
AI & ML
LLM Performance Testing | K6 + Grafana + InfluxDB | A tiny toolkit for load testing and benchmarking OpenAI-like inference endpoints using K6 + Grafana + InfluxDB
@0xnyn6d ago
AI & ML
Benchmark OpenAI-compatible AI endpoints and AI Accelerators in a reproducible structured way
@flexaihq6d ago
AI & ML
Distributed Inference with vLLM
@KempnerInstitute6d ago
DevTools
vLLM tool parser for Qwen2.5-Coder models using <tools> tag format.
@hanXen6d ago
AI & ML
FlashAttention-style custom attention backend for vLLM on AMD MI50/MI60/Radeon VII (gfx906). Downstream fork of mixa3607/ML-gfx906 with replacement HIP kernels and a vllm.general_plugins entry point.
@nick413-bit6d ago
AI & ML
Source-code level analysis of LLM RL training infra: async RL, weight sync, FP8, MoE routing | LLM RL
@zpqiu6d ago
DevTools
Hands-on GPU/HPC infrastructure operations: K8s GPU scheduling, HAMi sharing, Slurm, observability & vLLM inference. Learn it free on a laptop; validate on one cheap GPU.
@ld-singh6d ago
DevTools
npm pacakge to Craft files into Markdown with ease
@jparkerweb6d ago
DevTools
General Information, model certifications, and benchmarks for nm-vllm enterprise distributions
@neuralmagic6d ago
AI & ML
Extend LLM context windows beyond GPU memory limits with disk-backed KV cache.
@tirdyhouse6d ago
AI & ML
Terraform setup for deploying a private coding LLM on Vast.ai with vLLM, Qwen3 Coder, and OpenCode.
@anqorithm6d ago
AI & ML
LLM inference engine built from scratch in C++. No PyTorch, no frameworks.
@Anirudh1712026d ago
DevTools
The official implementation of CVPR Workshop 2025 paper: Window Token Concatenation for Efficient Visual Large Language Models.
@JackYFL6d ago
AI & ML
Ask Poddy: Run Open Source LLMs and Embeddings as OpenAI-Compatible Serverless Endpoints (Tutorial)
@blib-la6d ago
AI & ML
agentsculptor is an experimental AI-powered development agent designed to analyze, refactor, and extend Python projects automatically. It uses an OpenAI-like planner–executor loop on top of a vLLM backend, combining project context analysis, structured tool calls, and iterative refinement. It has only been tested with gpt-oss-120b via vLLM.
@Perpetue2376d ago
AI & ML
AI-Inference-Managed-by-AI: Go binary for managing AI inference on edge devices
@Approaching-AI6d ago
AI & ML
Multi-model LLM serving for NVIDIA DGX Spark with vLLM, web UI, and tool calling
@dataforgex6d ago
AI & ML
Tiny, stateless Go router that dispatches OpenAI-compatible requests to single-model vLLM and sglang backends with zero external dependencies
@mirkolenz6d ago
DevTools
Docker Compose setup that integrates vLLM, Open WebUI and NGINX (for SSL termination).
@marib006d ago
AI & ML
Honeycomb Lab — hex map + OpenAI gateway control plane for a home AI fleet
@joeynyc6d ago
AI & ML
A slim, local-LLM-first AI coding assistant. TUI, Web UI, and desktop app. Runs air-gapped with Ollama, vLLM, or any OpenAI-compatible endpoint.
@bobbyjohnstx22h ago
AI & ML
Automated comic cataloging tool that identifies issues directly from cover images using a vision-language model, then cross-references results with the Grand Comics Database and the ComicVine API to generate structured, high-confidence collection data with minimal manual entry.
@boyobob6d ago
DevTools
Spark Pulse is a web control plane for spark-vllm-docker
@kharkevich-engineering-lab6d ago
Dashboards
Two-node DGX Spark/ASUS GX10 DeepSeek V4 Flash DSpark NVFP4 port for vLLM 0.25, with live dashboard and reproducible deployment.
@Anemll22h ago
AI & ML
A curated collection of NLP and LLM resources. Covers essential papers and blogs on Transformers, Reinforcement Learning (RLHF, DPO, GRPO), Mechanistic Interpretability, Scaling Laws, and MLSys.
@rraghavkaushik6d ago
AI & ML
Open-source, self-hosted LLM chat application, featuring local-first data storage and real-time streaming responses.
@GHuyHuynh6d ago
AI & ML
AI-,. XTTS v2, real-time (Vosk/Whisper) offline LLM. - (Vue 3), Telegram-,, fine-tuning pipeline. Self-hosted,,.
@ShaerWare6d ago
AI & ML
A voice assistant with local LLM as a backend
@Saga91036d ago
AI & ML
Perl Framework for AI - Langertha - the viking of AI
@Getty6d ago
AI & ML
Repository for Multililngual Generation, RAG evaluations, and surrogate judge training for Arena RAG leaderboard (NAACL'25)
@vectara6d ago
From the blog

Latest vLLM guides & deep-dives

All articles →
Video archive

vLLM talks, tutorials & deep-dives

All videos →
$ made-with-vllm
rssllmmadewithwhat

Ask MadeWithWhat

AI answers may contain mistakes — please double-check important details.