// 333 open-source projects indexed

$ Discover the best projects built with vLLM

A hand-curated, daily-updated gallery of open-source projects. Browse the ecosystem, ranked by GitHub stars.
>
sort --stars
#all#aiml279#devtools44#docs4#dashboards3#blogs1
// 333 results
AI & ML
Open-source speech recognition toolkit for training, inference, streaming ASR, VAD, punctuation, speaker diarization pipelines, and OpenAI-compatible/MCP serving.
@modelscope4w ago
AI & ML
Welcome to the Llama Cookbook! This is your go to guide for Building with Llama: Getting started with Inference, Fine-Tuning, RAG. We also show you how to solve end to end problems using Llama model family and using them on various provider services
@meta-llama4mo ago
Blogs
Source Code Deep Dives, System Design & Engineering Blogs | Halfrost-Field :、
@halfrost4w ago
AI & ML
Comprehensive open-source library of AI research and engineering skills for any AI model. Package the skills and your claude code/codex/gemini agent will be an AI research agent with full horsepower. Maintained by Orchestra Research.
@Orchestra-Research3mo ago
AI & ML
LMCache: Supercharge Your LLM with the Fastest KV Cache Layer
@LMCache4w ago
AI & ML
An Easy-to-use, Scalable and High-performance Agentic RL Framework based on Ray (PPO & DAPO & REINFORCE++ & VLM & TIS & vLLM & Ray & Async RL)
@OpenRLHF7w ago
AI & ML
Swap GPT for any LLM by changing a single line of code. Xinference lets you run open-source, speech, and multimodal models on cloud, on-prem, or your laptop — all through one unified, production-ready inference API.
@xorbitsai4w ago
AI & ML
A Datacenter Scale Distributed Inference Serving Framework
@ai-dynamo4w ago
AI & ML
Mooncake is the serving platform for Kimi, a leading LLM service provided by Moonshot AI.
@kvcache-ai4w ago
AI & ML
Standardized Distributed Generative and Predictive AI Inference Platform for Scalable, Multi-Framework Deployment on Kubernetes
@kserve4w ago
AI & ML
A Low-Code MCP Framework for Building Complex and Innovative RAG Pipelines
@OpenBMB4w ago
AI & ML
A GPU cluster manager for high-performance AI model serving (vLLM, SGLang) and on-demand SSH-accessible GPU instances.
@gpustack4w ago
AI & ML
Reliable model swapping for any local OpenAI/Anthropic compatible server - llama.cpp, vllm, etc
@mostlygeek4w ago
AI & ML
Structured data extraction, instruction calling and agentic workflows with ML, LLM and Vision LLM
@katanaml4w ago
AI & ML
A programmable Mixture-of-Models router for heterogeneous LLM inference
@vllm-project4w ago
AI & ML
A course of learning LLM inference serving on Apple Silicon for systems engineers: build a tiny vLLM + Qwen.
@skyzh4w ago
AI & ML
Easiest and laziest way for building multi-agent LLMs applications.
@LazyAGI4w ago
AI & ML
Cascading runtime for AI agents. Optimize cost, latency, quality, and policy decisions inside the agent loop.
@lemony-ai2mo ago
AI & ML
High-performance Inference and Deployment Toolkit for LLMs and VLMs based on PaddlePaddle
@PaddlePaddle4w ago
AI & ML
The fastest local AI engine for Apple Silicon. 4.2x faster than Ollama, 0.08s cached TTFT, 100% tool calling. 17 tool parsers, prompt cache, reasoning separation, cloud routing. Drop-in OpenAI replacement. Works with Claude Code, Cursor, Aider.
@raullenchai4w ago
AI & ML
RamaLama is an open-source developer tool that simplifies the local serving of AI models from any source and facilitates their use for inference in production, all through the familiar language of containers.
@containers4w ago
AI & ML
Community maintained hardware plugin for vLLM on Ascend
@vllm-project4w ago
AI & ML
Control panel for VLLM, Sglang, llama.cpp, exllamav3
@sybil-solutions4w ago
AI & ML
A SOTA quantization algorithm for high-accuracy low-bit LLM inference, seamlessly optimized for CPU/XPU/CUDA, with multi-datatype support and full compatibility with vLLM, SGLang, and Transformers.
@intel4w ago
AI & ML
OpenAI and Anthropic compatible server for Apple Silicon. Run LLMs and vision-language models (Llama, Qwen-VL, LLaVA) with continuous batching, MCP tool calling, and multimodal support. Native MLX backend, 400+ tok/s. Works with Claude Code.
@waybarrios4w ago
AI & ML
Open Source Continuous Inference Benchmark Research Platform — Kimi K3 2.8T, MiniMax M3, DeepSeekv4, GLM5 - GB200 NVL72 vs MI355X vs B200 vs GB300 NVL72 & soon™ TPUv6e/v7/Trainium2/3 | — Kimi K2.7-Code、MiniMax M3、DeepSeekv4、GLM5 - GB200 NVL72 vs MI355X vs B200 vs GB300 NVL72,™ TPUv6e/v7/Trainium2/3
@SemiAnalysisAI4w ago
AI & ML
AI Inference Operator for Kubernetes. The easiest way to serve ML models in production. Supports VLMs, LLMs, embeddings, and speech-to-text.
@kubeai-project5w ago
AI & ML
LLM model quantization (compression) toolkit with HW acceleration support for Nvidia, AMD, Intel GPU and Intel/AMD/Apple CPU via HF, vLLM, and SGLang.
@ModelCloud4w ago
AI & ML
Enterprise-grade API gateway that helps you monitor and impose cost or rate limits per API key. Get fine-grained access control and monitoring per user, application, or environment. Supports OpenAI, Azure OpenAI, Anthropic, vLLM, and open-source LLMs.
@bricks-cloud1y ago
AI & ML
Virtualized Elastic KV Cache for Dynamic GPU Sharing and Beyond
@ovg-project4w ago
AI & ML
Evaluate your LLM's response with Prometheus and GPT4
@prometheus-eval1y ago
AI & ML
Build your own high performance LLM inference engine in C++ and CUDA - a smaller version of vLLM
@jmaczan4w ago
AI & ML
LLM notes, including model inference, transformer model structure, and llm framework code analysis notes.
@harleyszhang5w ago
DevTools
UniRL is a Framework for Unified Multimodal Model Reinforcement Learning
@Tencent-Hunyuan4w ago
AI & ML
Make Discord your LLM frontend - Supports any OpenAI compatible API (OpenRouter, Ollama and more)
@jakobdylanc2mo ago
AI & ML
A Python library powered by Language Models (LLMs) for conversational data discovery and analysis.
@pgalko3mo ago
AI & ML
An local, offline (after initial setup), portable OCR software that can process images and PDF files, using DeepSeek-OCR-2 AI (running directly on your machine).
@th1nhhdk2mo ago
DevTools
[EMNLP 2024 & AAAI 2026] A powerful toolkit for compressing large models including LLMs, VLMs, and video generative models.
@ModelTC4mo ago
Docs
Multimodal RL training framework for diffusion & omni models
@verl-project4w ago
AI & ML
Accurate, large-scale, and extensible simulator for LLM inference Systems
@microsoft1y ago
AI & ML
EfficientSAM3 compresses SAM3 into lightweight, edge-friendly models via progressive knowledge distillation for fast promptable concept segmentation and tracking.
@SimonZeng71082mo ago
AI & ML
Pure Rust + CUDA LLM inference engine — no PyTorch, OpenAI-compatible, serves Qwen3 to Kimi-K2
@openinfer-project4w ago
AI & ML
AI Bank Statement Document Automation By LLM model and Personal Finanical Analysis
@johnsonhk884w ago
DevTools
SparkTTS、OrpheusTTS,。
@HuiResearch1y ago
AI & ML
Fine-tune LLMs from one YAML. Layer streaming trains an 8B model on a 4 GB laptop GPU.
@MakazhanAlpamys3w ago
AI & ML
[CVPR 2025] RoboBrain: A Unified Brain Model for Robotic Manipulation from Abstract to Concrete. Official Repository.
@FlagOpen11mo ago
AI & ML
Crater is a cloud-native AI training & inference platform.
@raids-lab4w ago
AI & ML
RESTai is an AIaaS (AI as a Service) open-source platform. Supports many public and local LLM suported by Ollama/vLLM/etc. Precise embeddings usage, tuning, analytics etc. Built-in image/audio generation with dynamic loading generators. Live chat deployment. Built-in block based graphical language. Prompt versioning and much more...
@apocas3w ago
DevTools
A command-line interface tool for serving LLM using vLLM.
@Chen-zexi7mo ago
AI & ML
A modern web interface for managing and interacting with vLLM servers (www.github.com/vllm-project/vllm). Supports both GPU and CPU modes, with special optimizations for macOS Apple Silicon and enterprise deployment on OpenShift/Kubernetes.
@micytao5mo ago
AI & ML
Open Model Engine (OME) — Kubernetes operator for LLM serving, GPU scheduling, and model lifecycle management. Works with SGLang, vLLM, TensorRT-LLM, and Triton
@ome-projects4w ago
AI & ML
Self-hosted AI agent OS. Your memory, chat, agents, and files stay on hardware you own, offline by default, cloud by choice. Offline AI memory (taOSmd), self-hosted multi-framework group chat, a full web desktop + app store, and auto-clustering across the consumer hardware you already have (Orange/Raspberry Pi, Mac mini, gaming PC).
@jaylfc4w ago
AI & ML
The open-source, agent-first creative workspace.
@nodetool-ai3w ago
AI & ML
The Runpod worker template for serving our large language model endpoints. Powered by vLLM.
@runpod-workers4w ago
AI & ML
KVarN is a native vLLM KV-cache quantization backend for your agents: 3-5x more context, throughput above FP16, and FP16-level accuracy. Calibration-free, one flag.
@huawei-csl2mo ago
AI & ML
Engine-agnostic LLM gateway in Rust. Full OpenAI & Anthropic API compatibility across vLLM, TRT-LLM, TokenSpeed, SGLang, OpenAI, Gemini & more. Industry-first gRPC pipeline, KV cache-aware routing, chat history, tokenization caching, Responses API, embeddings, WASM plugins, MCP, and multi-tenant auth.
@smg-project4w ago
AI & ML
Ready-to-use DeepSeek-OCR Web UI | Modern Interface | 7 Recognition Modes | Batch Processing | Real-time Logging | Fully Responsive
@neosun1006mo ago
AI & ML
Fully uncensored, capability-enhanced abliteration of Qwen3.6-27B. NVFP4 + z-lab DFlash speculative decoding (n=12) on the unified ghcr.io/aeon-7/aeon-vllm-ultimate:latest container, tuned for long-context draft acceptance on DGX Spark. 6 HF variants (BF16/NVFP4/MTP/MTP-XS), docker-compose, and QuickStart.
@AEON-72mo ago
AI & ML
sparkrun - launch, manage, and stop LLM inference workloads on NVIDIA DGX Spark systems
@spark-arena4w ago
AI & ML
TopicGPT: A Prompt-Based Framework for Topic Modeling [NAACL'24]
@chtmp2233mo ago
AI & ML
LvLLM is a special NUMA extension of vllm that makes full use of CPU and memory resources, reduces GPU memory requirements, and features an efficient GPU parallel and NUMA parallel architecture, supporting hybrid inference for MOE large models.
@guqiong964w ago
AI & ML
Python package for LLM compression
@FujitsuResearch4w ago
AI & ML
Low latency JSON generation using LLMs
@varunshenoy2y ago
AI & ML
Setup and run a local LLM and Chatbot using consumer grade hardware.
@jasonacox9mo ago
AI & ML
Versatile Almost Local, Eventually Reasonable Assistant
@vakovalskii4mo ago
AI & ML
Persist and reuse KV Cache to speedup your LLM.
@ModelEngine-Group4w ago
AI & ML
Easy, advanced inference platform for large language models on Kubernetes. Star to support our work!
@InftyAI10mo ago
Docs
Implementation for FP8/INT8 Rollout for RL training without performence drop.
@yaof2010mo ago
DevTools
Qwen3.5-122B-A10B on DGX Spark: 28.3 → 51 tok/s (+80%)
@albond3mo ago
DevTools
AI-powered offensive security testing using autonomous agents, directly in your terminal.
@pensarai4w ago
AI & ML
Blazing-fast LLM inference in pure Rust. No PyTorch and Python runtime.
@guoqingbao4w ago
AI & ML
The production framework for Predictive and Generative AI. Serve any model as an API in one line, with OpenAI/Anthropic/Ollama-compatible endpoints, a built-in chat UI, and native MCP.
@vortico5w ago
DevTools
ML software (llama.cpp, ComfyUI, vLLM) builds for AMD gfx906 GPUs, e.g. Radeon VII / MI50 / MI60
@mixa36074w ago
AI & ML
A 10-week, 30-minutes-a-day roadmap for LLM inference serving and optimization. vLLM, SGLang, quantization, speculative decoding, benchmarking.
@patchy6314w ago
AI & ML
CacheRoute is an innovative LLM scheduling scheme dedicated to enabling flexible KV cache reuse across LLM systems, improving task performance and system efficiency.
@AstraNetLab4w ago
AI & ML
High-performance lightweight proxy and load balancer for LLM infrastructure. Intelligent routing, automatic failover and unified model discovery across local and remote inference backends.
@thushan3w ago
AI & ML
LLM ,。
@lework9mo ago
AI & ML
[ICML-2026] Official implementation of "SEAgent: Self-Evolving Computer Use Agent with Autonomous Learning from Experience"
@SunzeY1y ago
AI & ML
An Open-source, self-hosted AI model hub with Hugging Face compatibility, accelerating vLLM/SGLang performance.
@matrixhub-ai4w ago
AI & ML
A CPU Realtime VLM in 500M. Surpassed Moondream2 and SmolVLM. Training from scratch with ease.
@lucasjinreal1y ago
AI & ML
Home Assistant LLM integration for local OpenAI-compatible services (llamacpp, vllm, etc)
@skye-harris4w ago
AI & ML
gpt_serverLLMs、Embedding、Reranker、ASR、TTS、、。
@shell-nlp4mo ago
AI & ML
One-click Qwen3.6-27B inference on Windows. 158 tok/s on RTX 5090, 72 tok/s on RTX 3090. Native, no WSL, no Docker, no telemetry.
@devnen4mo ago
AI & ML
A PyTorch native library for training speculative decoding models
@lightseekorg4w ago
AI & ML
a fun and educational take on vLLM
@ovshake7mo ago
AI & ML
Simple LLM service identification - translate IP:Port to Ollama, vLLM, LiteLLM, or 60+ other AI services in seconds
@praetorian-inc4w ago
AI & ML
A general-purpose API load testing platform that supports LLM services and business HTTP interfaces, enabling one-click performance testing, result comparison, and AI-powered intelligent analysis and summarization. API , HTTP ,、 AI
@MigoXLab8w ago
AI & ML
Open-source tools for training and evaluating Vision Language Models for OCR
@Roots-Automation2mo ago
Docs
vLLM Documentation in Chinese Simplified / vLLM
@hyperai6mo ago
AI & ML
Fully-featured, beautiful web interface for vLLM - built with NextJS.
@yoziru4w ago
AI & ML
Kubernetes operator for self-hosted LLM inference across a heterogeneous GPU fleet: NVIDIA CUDA, AMD Vulkan, and Apple Silicon Metal. Runtimes: llama.cpp, vLLM, TGI, mlx-server. Multi-GPU sharding, model caching, OpenAI-compatible endpoints. Apache-2.0, run across homelab and on-prem fleets, actively developed.
@defilantech4w ago
AI & ML
[ICCV2025] Referring any person or objects given a natural language description. Code base for RexSeek and HumanRef Benchmark
@IDEA-Research11mo ago
AI & ML
High-performance KV cache storage for LLM inference — GPU offloading, SSD caching, and cross-node sharing via RDMA. Works with vLLM and SGLang.
@novitalabs4w ago
AI & ML
Deep Learning Deployment Framework: Supports tf/torch/trt/trtllm/vllm and other NN frameworks. Support dynamic batching, and streaming modes. It is dual-language compatible with Python and C++, offering scalability, extensibility, and high performance. It helps users quickly deploy models and provide services through HTTP/RPC interfaces.
@NetEase-Media1y ago
AI & ML
Booster - open accelerator for LLM models. Better inference and debugging for AI hackers
@gotzmann2y ago
AI & ML
Probing Scientific General Intelligence of LLMs with Scientist-Aligned Workflows
@InternScience3mo ago
AI & ML
Automated system for LLM evaluation via agents. Doc as below:
@OpenDCAI5w ago
AI & ML
Cognithor · Agent OS: Local-first autonomous agent operating system. 19 LLM providers, 18 channels, 145 MCP tools, 6-tier memory, Agent Packs marketplace, zero telemetry. Python 3.12+, Apache 2.0.
@Alex8791-cyber4mo ago
AI & ML
GLM-5.2-NVFP4-REAP-469B serving on SM120 (4× RTX PRO 6000 Blackwell) — one-command vLLM launch recipe, 250K context, DeepSeek Sparse Attention + MTP speculative decode
@0xSero3mo ago
AI & ML
Never give AI companies your secrets! A local LLM-based privacy filter for LLM users. Seamless integration with your existing AI tools as a Python library / OpenAI SDK replacement / API Gatetway / Web Server.
@cxumol8mo ago
AI & ML
DocMind AI is a powerful, open-source Streamlit application leveraging LlamaIndex, LangGraph, and local Large Language Models (LLMs) via Ollama, LMStudio, llama.cpp, or vLLM for advanced document analysis. Analyze, summarize, and extract insights from a wide array of file formats, securely and privately, all offline.
@BjornMelin7w ago
DevTools
vLLM-5090: Docker Container for RTX 5090 + OpenCode
@BoltzmannEntropy9mo ago
AI & ML
Unified management and routing for llama.cpp, MLX and vLLM models with web dashboard.
@lordmathis5w ago
AI & ML
SNDR Core Engine (Genesis) — vLLM runtime patch-overlay for Qwen3.6 + Gemma4 on consumer NVIDIA (Ampere sm_86, 2× A5000/3090). Qwen3.6-35B-A3B FP8 ~240 tok/s, 27B-int4 hybrid GDN+Mamba, Gemma4 26B/31B AWQ, 256K ctx. 321 patches: TurboQuant k8v4 KV, MTP/DFlash spec-decode, FULL cudagraph, hybrid GDN. vLLM pin dev424 + Control Center GUI.
@Sandermage4w ago
AI & ML
Privacy-first LLM proxy and AI gateway - load balancing, multi-provider routing, API key management, usage tracking, rate limiting. Self-hosted. Zero knowledge of your prompts.
@voidmind-io6w ago
AI & ML
A project worth exploring.
@GaeaRuiW3mo ago
AI & ML
Your AI intranet: network the computers you already own for inference and training.
@autonomous-ai4w ago
AI & ML
A High-Performance LLM Inference Engine with vLLM-Style Continuous Batching
@lumia4318mo ago
AI & ML
KEDA External gRPC Scaler for GPU workloads - native NVML metrics via DaemonSet, no Prometheus required
@pmady5w ago
AI & ML
Convert PowerPoint files into semantically rich text using vision language models
@ALucek10mo ago
DevTools
Data Center and Client workload and software optimizations for Intel hardware.
@intel4w ago
AI & ML
Efficient LLM inference on Slurm clusters.
@VectorInstitute4w ago
AI & ML
One-command vLLM installation for NVIDIA DGX Spark with Blackwell GB10 GPUs (sm_121 architecture)
@eelbaz10mo ago
DevTools
Local diagnostic CLI for NVIDIA DGX Spark (GB10). Detects power caps, UMA pressure, thermal risk, CUDA 13/SM_121 wheel mismatches, Docker issues, and validates vLLM/Ollama/llama.cpp/SGLang recipes.
@joeynyc8w ago
AI & ML
Kubernetes-native platform for deploying and managing AI inference across multiple providers
@ai-runway4w ago
AI & ML
Practical local LLM recipes and benchmarks for RTX 5060 Ti setups
@5p00kyy8w ago
AI & ML
llm-inference is a platform for publishing and managing llm inference, providing a wide range of out-of-the-box features for model deployment, such as UI, RESTful API, auto-scaling, computing resource management, monitoring, and more.
@OpenCSGs2y ago
AI & ML
Nano vLLM with vLLM v1's request scheduling strategy and chunked prefill
@slwang-ustc7mo ago
AI & ML
FlowSteer: agents designing agentic workflows via reinforced progressive canvas editing.
@beita69694mo ago
AI & ML
Extensible generative AI platform on Kubernetes with OpenAI-compatible APIs.
@llmariner4mo ago
AI & ML
A comprehensive toolkit for deploying production-ready Generative AI infrastructure on Amazon EKS. Includes pre-configured components for: AI Gateway (LiteLLM) LLM Serving (vLLM, SGLang, Ollama) Vector Databases, Embedding Models (TEI) Observability (Langfuse, Phoenix) etc. Fast-track your GenAI deployment with Kubernetes
@aws-samples4w ago
AI & ML
This repository contains hands-on projects, code examples, and deployment workflows. Explore multi-agent systems, LangChain, LangGraph, AutoGen, CrewAI, RAG, MCP, automation with n8n, and scalable agent deployment using Docker, AWS, and BentoML.
@MDalamin53mo ago
AI & ML
Granite Switch — Build AI models like you build software
@generative-computing4w ago
AI & ML
Qwen2+SFT+DPO, SFTTrainer/DPOTrainer/TRPOTrainer,,(neo4j, milvus, LDA, )。, vllm, vllm API embedder + Reranker RAG 。 MDAgents , vllm api 。
@NJUxlj4mo ago
AI & ML
Enterprise-grade LLM automated deployment tool that makes AI servers truly "plug-and-play".
@EM-GeekLab5mo ago
AI & ML
Zero trust LLM gateway. OpenAI-compatible proxy with semantic routing and load balancing across OpenAI, Anthropic, Ollama, vLLM, and any compatible backend. Identity-based access, virtual API keys, and end-to-end encryption via OpenZiti
@openziti5w ago
Dashboards
Real-time hardware and LLM inference monitoring — GPU, CPU, memory, and vLLM metrics streamed to a dashboard.
@niklasfrick5w ago
AI & ML
A unified evaluation toolkit and leaderboard for rigorously assessing the scientific intelligence of large language and vision–language models across the full research workflow.
@InternScience3mo ago
AI & ML
The open-core AI workbench — notebooks, agents, RAG, voice, and images across any model: OpenAI, Anthropic, Google, xAI, or local via Ollama/vLLM. BSL 1.1, auto-converting to Apache-2.0 on a two-year clock. Your AI keeps running when theirs doesn't.
@Bike4Mind3w ago
AI & ML
Agent and LLM API performance monitoring via network packet probe. Measures performance of OpenClaw, Claude, Codex, DeepAgents and more — deployed on the provider side, no SDK changes required.
@Netis2mo ago
AI & ML
A High-Efficiency System of Large Language Model Based Search Agents
@tiannuo-yang1y ago
AI & ML
A simple service that integrates vLLM with Ray Serve for fast and scalable LLM serving.
@asprenger2y ago
AI & ML
First open-source implementation of Google TurboQuant (ICLR 2026) -- near-optimal KV cache compression for LLM inference. 5x compression with near-zero quality loss.
@OnlyTerp3mo ago
AI & ML
Reliable and Efficient Semantic Prompt Caching with vCache
@vcache-project5w ago
AI & ML
Production inference for encoder models - ColBERT, GLiNER, ColPali, embeddings etc. - as vLLM plugins for online and in-process deployment
@latenceainew2mo ago
AI & ML
[ICML 2026] Decoding Tree Sketching (DTS): a training-free & model agonistic & plug-in framework for LLM parallel reasoning.
@ZichengXu4mo ago
AI & ML
This project aims to provide a high effective KV cache manage framework for llm inference and improve memory utilization and inference speed.
@TheToughCrane4mo ago
AI & ML
The htop for LLM inference see exactly where every GB of VRAM goes and get measured quantization savings.
@helasaoudi3w ago
AI & ML
EmbodiedAgents is a fully-loaded ROS2 based framework for creating interactive physical agents that can understand, remember, and act upon contextual information from their environment.
@automatika-robotics3w ago
Dashboards
An intelligent tuner for vLLM that automatically monitors GPU metrics, uses Bayesian optimization to tune parameters
@jranaraki6mo ago
AI & ML
PDF Parsing Tool: GOT's vLLM acceleration implementation, MinerU for layout recognition, and GOT for table formula parsing.
@liunian-Jay1y ago
AI & ML
Automated Deep Research with LLMs, web search, paper parsing, and didactic summarization.
@protonspy1y ago
AI & ML
Community registry of LLM serving-path traps that produce confidently wrong measurements: templates, tool parsers, reasoning fields, quant kernel paths, CUDA toolchains, KV allocation, eval harnesses, versioning. Symptom-first, with the check that catches each.
@Blackwellboy4w ago
AI & ML
Real-time speech-to-text WebSocket server with pluggable ASR backends, energy-based VAD, streaming partial results, and Prometheus observability.
@victoryangzhijie5w ago
AI & ML
LLM-Inference-Bench
@argonne-lcf1y ago
AI & ML
A tool for benchmarking LLMs on Modal
@modal-labs1y ago
AI & ML
Turn documents into structured JSON with local-first document AI. Run 100% locally by default, with API, CLI, and Web UI.
@parsehawk4w ago
DevTools
Sardeenz is a proof-of-concept application that allows you to load more than one model on a given GPU. It allows you to add more and more models onto a GPU, until it is fully utilized.
@rh-aiservices-bu3mo ago
AI & ML
LLM-powered security log analyzer: detect threats & anomalies with zero regex — just declare a Pydantic schema. Real-time Telegram alerts, SIEM-ready with Elasticsearch/Kibana. Supports OpenAI, Ollama, vLLM.
@call5187w ago
AI & ML
An endpoint server for efficiently serving quantized open-source LLMs for code.
@wangcx182y ago
DevTools
A physics-grounded, cost-aware optimization loop for vLLM
@jungledesh4w ago
AI & ML
Modern AI chatbot supporting multiple LLMs. Switch between Gemini, Mistral, Llama, Claude and ChatGPT.
@intelligentnode1y ago
AI & ML
vLLM Router
@llm-semantic-router2y ago
AI & ML
High-Performance KV Cache Storage Engine on CXL Shared Memory for LLM Inference
@xcena-dev5w ago
AI & ML
OpenAI-compatible multilingual TTS server — Chatterbox on vLLM with real-time PCM audio streaming, low time-to-first-byte (~0.7 s), voice cloning, and 23 languages.
@wuxuedaifu3mo ago
Real-time
Real-time streaming TTS server for XTTS-v2 on vLLM — OpenAI-compatible API, ~0.5s TTFB, Docker
@wuxuedaifu8w ago
AI & ML
The open-source AI platform for enterprises that can't send data to the cloud. OpenAI-compatible API, full management dashboard, zero data egress.
@xinity-ai4w ago
AI & ML
DFlash vLLM for DGX Spark — Plug & Play Block-Diffusion Speculative Decoding
@AEON-72mo ago
AI & ML
DGX Spark / GB10 vLLM Docker stack for large-model serving, presets, patches, and validation notes.
@bjk1106w ago
AI & ML
[ICPP'25] TD-Pipe: Temporally-Disaggregated Pipeline Parallelism Architecture for High-Throughput LLM Inference
@MLSysU8mo ago
AI & ML
vLLM Qwen 3.6-27B (AWQ-INT4) + DFlash speculative decoding on AMD Strix Halo (gfx1151 iGPU, 128 GB UMA, ROCm 7.13). 24.8 t/s single-stream, vision, tool calling, 256K context, OpenAI-compatible, Docker. Matches DGX Spark FP8+DFlash+MTP at a third of the cost. No CUDA.
@hec-ovi4mo ago
AI & ML
Arks is a cloud-native inference framework running on Kubernetes
@scitix4mo ago
AI & ML
Open source Enterprise AI Gateway for API, MCP and agent, and AI model traffic. One Apache-2.0 binary: 72 native providers behind an OpenAI-compatible API, or serve vLLM and llama.cpp on your own GPUs. Keys, budgets, guardrails, semantic cache, WAF.
@soapbucket4w ago
AI & ML
Batch Deployment for Document Parsing with AWS Batch & Qwen-2.5-VL
@jeremyarancio1y ago
AI & ML
TurboQuant KV cache compression plugin for vLLM — asymmetric K/V, 8 models validated, consumer GPUs
@Alberto-Codes5mo ago
AI & ML
A high-performance RDMA distributed file system for fast LLM Inference and GPU Training.
@blackbird-io7mo ago
AI & ML
Native macOS menu bar app for realtime dictation with optional LLM polishing. Connects to any OpenAI Realtime-compatible backend — fully local on Apple Silicon with voxmlx + mlx-swift-lm.
@T0mSIlver7w ago
AI & ML
a community oriented 1:1, vLLM-alike (Continuous batching, paged KV) engine in C++ with additional features
@mudler4w ago
AI & ML
ElasticMM: Elastic and Efficient MLLM Serving System
@hpdps-group4mo ago
AI & ML
Local-first CLI for benchmarking LLMs on real hardware — quality, speed, reliability, and a real multi-turn agent loop.
@outsourc-e3mo ago
AI & ML
Android keyboard with local AI (Ollama, Whisper, MCP) or cloud (Gemini, Groq, OpenAI)
@SvReenen4mo ago
DevTools
useful Gentoo overlay Curated ebuilds, AI, tools & science
@istitov4w ago
AI & ML
Bleeding edge vLLM Docker image for the NVIDIA DGX Spark (GB10 / sm_121a).
@timothystewart65w ago
AI & ML
Private LLM/RAG platform in one command for NVIDIA DGX Spark / GB10 (arm64). Validated on real hardware.
@botAGI5w ago
Docs
Synthetic medical VQA pipeline: 119K images annotated by frontier VLMs, cross-validated at 93% agreement, fine-tuned on 3 model families (2-3B params)
@openmed-labs5mo ago
AI & ML
Carbon Limiting Auto Tuning for Kubernetes
@Climatik-Project6mo ago
DevTools
Code for building self-expanding knowledge graphs with Outlines, vLLM, neo4j, and Modal.
@just-cameron1y ago
AI & ML
Hit your limit? Need privacy? Just swap the model, everything else stays
@luongnv895w ago
AI & ML
Open-source, self-hostable social listening.
@obris-dev6w ago
AI & ML
Self-hosted, OpenAI-compatible inference for the agentic era: reasoning LLMs, universal tool calling, and the Responses API alongside embeddings, speech, and image models — many models sharing your GPUs, one gateway. Powered by Ray Serve.
@modelship-ai4w ago
Dashboards
Two-node DGX Spark/ASUS GX10 DeepSeek V4 Flash DSpark NVFP4 port for vLLM 0.25, with live dashboard and reproducible deployment.
@Anemll6w ago
DevTools
GLM-5.2 (744B/40B MoE) on a 4× DGX Spark / GB10 (sm_121) cluster: portable Triton sparse-MLA kernels, a data-free expert prune, MTP draft, and a one-script bootstrap.
@CosmicRaisins5w ago
AI & ML
Native Windows vLLM 0.26.0: CPython 3.13, CUDA 12.8, SM 7.5/8.6/8.9/12.0 for RTX 20/30/40/50, OpenAI-compatible serving, Triton/FlashAttention, 10 KV-cache formats, Multi-TurboQuant, and experimental CPU/RAM/NVMe prompt-KV offload. No WSL or Docker.
@aivrar4w ago
AI & ML
From-scratch C++/CUDA inference engine for the NVIDIA RTX 5090 (sm_120a) — the best single-GPU backend for agentic AI: tool calling, long-context loops, reasoning and concurrent sub-agents on top of the fastest single-stream decode on the 5090 (beats llama.cpp, at-or-ahead of vLLM on NVFP4). 100% written by Claude Code.
@kekzl4w ago
DevTools
Benchmarking Open-Ended Inference Optimization by AI Agents
@aisa-group2mo ago
AI & ML
An 18 notebook course that isolates and measures each component of agentic loop engineering on real, industry standard software datasets.
@FareedKhan-dev2mo ago
AI & ML
AIfred-Intelligence — self-hosted Multi-Agent Assistant with Debate Modes (Symposion/Tribunal), Voice (STT + Streaming-TTS), RAG with Long-Term Memory, Web Research and Tool-Calling. Reachable via Web-UI or Telegram/Discord/Email/EPIM. Multi-Backend: llama-swap, Ollama, vLLM, TabbyAPI.
@Peuqui3w ago
AI & ML
Lightweight proxy for LLM
@yatesdr5w ago
DevTools
vllm
@gameofdimension2y ago
DevTools
Serve Llama 3.3 70B (with AWQ quantization) using vLLM and deploy it on BentoCloud.
@kingabzpro1y ago
AI & ML
AI Coding Assistant CLI for offline enterprise environments - Local LLM platform with Plan & Execute architecture, Supervised Mode, and auto-update system
@HanSyngha4mo ago
AI & ML
Powerful no-code LLM fine-tuner: upload data → train → deploy in minutes. Unsloth 2-5× acceleration · QLoRA/DPO/RLHF/PPO/ORPO · Reward Model training · GGUF export · vLLM inference · BLEU/ROUGE/BERTScore · full CLI · Heretic Mode to unlock full model potential
@Yog-Sotho4w ago
AI & ML
A REST API for vLLM, production ready
@France-Travail11mo ago
DevTools
Policy-driven seamless lazy loading
@cloudpilot-ai5w ago
AI & ML
[CVPR2025] Official Repository for IMMUNE: Improving Safety Against Jailbreaks in Multi-modal LLMs via Inference-Time Alignment
@itsvaibhav011y ago
AI & ML
a simple lightweight large language model pipeline framework.
@sherlockchou861y ago
AI & ML
A web-based memory usage and performance calculator for Huggingface GGUF models
@gdevenyi6w ago
AI & ML
Serve GLM-5.2 469B (REAP-pruned, NVFP4) across 3× NVIDIA DGX Spark with vLLM pipeline parallelism — 256K context, production-ready config and patches
@bird7w ago
AI & ML
An open-source, AI-driven Social Media Digital Twin powered by LLMs (Ollama, vLLM) for Computational Social Science simulations, network analysis, and human-agent interaction.
@YSocialTwin3w ago
AI & ML
Dockerized LLM inference server with constrained output (JSON mode), built on top of vLLM and outlines. Faster, cheaper and without rate limits. Compare the quality and latency to your current LLM API provider.
@phospho-app2y ago
AI & ML
Experimental RAG playground for exploring retrieval quality, corpus construction, and filter-chain design. Features configurable ranking and filtering pipelines, visual document grounding, chat interfaces, web search, Open WebUI integration, and rich debugging insights. Ollama and vLLM, running in Dev Containers or natively on Linux and Windows.
@HarinezumIgel4w ago
AI & ML
A hybrid router that uses Spot GPU instances to reduce costs and Serverless GPUs for making Cold Starts faster.
@Tandemn-Labs6mo ago
AI & ML
A "standard library" of Triton kernels.
@stackav-oss11mo ago
AI & ML
Official implementation for Text Generation Beyond Discrete Token Sampling
@EvanZhuang1y ago
AI & ML
Dual-engine (llama.cpp + vLLM) LLM benchmarking pipeline for GGUF & safetensors on NVIDIA GPUs — speed, quality, live dashboard, publishable cards.
@notwitcheer3w ago
AI & ML
Orpheus TTS Server with streaming support (TTFB ~160ms)
@taresh1811mo ago
AI & ML
Self-hosted meeting transcription portal — speech-to-text, speaker diarization, LLM-corrected transcripts, structured summaries and Word minutes, on your own GPUs. Flask + PostgreSQL, GDPR audit trail, distributed GPU topologies, docker
@Martossien3w ago
AI & ML
Python SDK for Agent Vector Protocol – transfer KV-cache between LLM agents instead of text
@VectorArc4mo ago
AI & ML
Patched native-Windows build of vLLM. Three Windows-specific fixes (CPU-relay for Gloo, Qwen3 reasoning parser, wildcard model name) on top of SystemPanic 0.19.0. Engine for devnen/qwen3.6-windows-server.
@devnen4mo ago
AI & ML
Unified KV-cache compression for LLM inference: 12 Python-native methods, Debian-tested isolated add-ons, Godzilla KVarN/TriAttention, exact Godzilla/Gigatoken profiles, CUDA weight sharing, and multi-GPU planning.
@aivrar4w ago
AI & ML
Rust SDK for building AI agents with local OpenAI-compatible servers (LMStudio, Ollama, llama.cpp, vLLM). Features streaming, tools, hooks, retry logic, and comprehensive examples.
@slb3503w ago
AI & ML
Probing the limitations of multimodal language models for chemistry and materials research
@lamalab-org7mo ago
AI & ML
Local AI workstation — discover, run, chat, benchmark, and generate images from open-weight models. DFlash/DDTree speculative decoding, TurboQuant & TriAttention cache compression strategies, MLX + llama.cpp + vLLM + MTPLX backends.
@cryptopoly5w ago
AI & ML
This project is the backend engine for a fully autonomous AI-powered call center. It integrates a large language model (LLM), speech recognition, and text-to-speech to manage real-time phone conversations via Asterisk.
@klikz-dev1y ago
AI & ML
Training-free KV cache compression via E8 lattice VQ. 2-bit KV that preserves retrieval (30/30 NIAH vs TurboQuant 0/30). Calibration-free, 9 architectures validated.
@jagmarques4w ago
AI & ML
Run the LLM Swarm Router on machines to distribute Local Ai to the Swarm - More Machines - MORE SPEED
@matthewdcage3w ago
AI & ML
, LLM
@zRzRzRzRzRzRzR2y ago
AI & ML
Say 'dreamier' and your ComfyUI workflow shifts — instantly, reversibly. An AI co-pilot for VFX artists: zero-LLM recipes (dreamier/sharper/faster), 133 MCP tools, EXR-aware vision, workflow.lock provenance, full undo, native sidebar, 5 swappable brains (Claude, GPT, Gemini, Ollama, Nemotron).
@JosephOIbrahim6w ago
AI & ML
An imperative command-line-interface for AI workload orchestration
@theoddden3w ago
AI & ML
Local-first desktop assistant for the whole job hunt — scrape listings, match with a local LLM, generate applications in your own voice, and track everything on a Kanban board.
@Keljian3w ago
AI & ML
NVFP4 inference on Blackwell GeForce (RTX 5090/5080/5070 Ti/RTX PRO 6000) — SM120 patches for vLLM + FlashInfer + CUTLASS. 175 tok/s on Qwen3.6-35B MoE.
@lna-lab4mo ago
AI & ML
Enterprise-grade platform to generate and execute Cypress, Playwright, WebdriverIO, and Appium end-to-end tests from natural language requirements. (Appium is experimental and requires external mobile infrastructure.)
@aiqualitylab3w ago
AI & ML
Open-source, self-hosted AI workspace for local and open-weight LLMs with vLLM, LiteLLM, autonomous agents, MCP tools, deep research, artifacts, and BYOK.
@openmake4w ago
DevTools
Run and manage both vLLM and llama.cpp model servers from one terminal UI + CLI — shared Docker Compose profiles, live tok/s, and cross-backend port/GPU conflict checks. NVIDIA + Linux.
@Changroro4w ago
DevTools
DeepSeek-V3, R1 671B on 8xH100 Throughput Benchmarks
@dzhsurf1y ago
AI & ML
Run the AEON Bench suite on your own hardware: verified HuggingFace pull → serve → benchmark (text · agentic ×3 harnesses · vision · audio · arena · perf) → ed25519-signed attested submit.
@AEON-74w ago
AI & ML
[ACL 2026] CoCoA: Collaborative Chain-of-Agents for Parametric-Retrieved Knowledge Synergy
@liunian-Jay11mo ago
AI & ML
OpenAI-compatible, vLLM-served OCR API for the Surya-OCR-2 model — multilingual document OCR (layout + text recognition) with request batching, a local CLI, and Docker packaging.
@wuxuedaifu2mo ago
AI & ML
Why is LLM inference slow — and how do you make it fast? A hands-on, first-principles course: roofline → KV cache → quantization → parallelism → vLLM/SGLang, with GPU labs on open models.
@jiahongsigma2mo ago
AI & ML
Low-Cost Cross-Domain Web Structured Information Extraction using specialized LoRA adapters.
@abdo-Mansour5w ago
AI & ML
Layered prefill changes the scheduling axis from tokens to layers and removes redundant MoE weight reloads while keeping decode stall free. The result is lower TTFT, lower end-to-end latency, and lower energy per token without hurting TBT stability.
@scale-snu6mo ago
AI & ML
RAG using LlamaIndex:Computer Network Q&A System powered by LlamaIndex | LlamaIndex - HyDE+ + vLLM +Ragas
@userHanlh6mo ago
AI & ML
A Scheduler for Batched LLM Inference
@iPieter11mo ago
AI & ML
ICE-PIXIU:A Cross-Language Financial Megamodeling Framework
@YY06491y ago
AI & ML
Comprehensive, scalable ML inference architecture using Amazon EKS, leveraging Graviton processors for cost-effective CPU-based inference and GPU instances for accelerated inference. Guidance provides a complete end-to-end platform for deploying LLMs with agentic AI capabilities, including RAG and MCP
@aws-solutions-library-samples6mo ago
AI & ML
Manalyzer: End-to-end Automated Meta-analysis with Multi-agent System
@black-yt3mo ago
AI & ML
Intelligent load balancer for distributed vLLM server clusters vLLM
@xerrors10mo ago
DevTools
[AAAI 2025 oral] Attribution Analysis Meets Model Editing: Advancing Knowledge Correction in Vision Language Models with VisEdit
@qizhou0001y ago
AI & ML
A lightweight post-training framework for LLMs and VLMs. 51 algorithms, 38 verified models. Scales with DeepSpeed, vLLM, and Ray.
@warlockee4w ago
AI & ML
Uncensored AI platform — chat, roleplay, image generation. Built on open-source models.
@Kyns-ai7w ago
AI & ML
Serve Poolside Laguna S 2.1 (NVFP4) on the NVIDIA DGX Spark (GB10) without hanging your box. Working stack, crash-safe configs, benchmarks, and the exact gotchas.
@sudoingX6w ago
AI & ML
Optimized vLLM setup for Qwen3.6-27B-FP8 on dual RTX PRO 6000 Blackwell (192 GB GDDR7, no NVLink); config, benchmark sweep results, and custom chat template with thinking mode off by default.
@theogravity4mo ago
AI & ML
Object-storage-native KV cache for LLM inference & RL. Cross-restart, cross-conversation, cross-engine via shared S3 bucket.
@Venkat28115w ago
DevTools
DGX Spark research and tests - containers, benchmarks, and investigation notes for running models on GB10 (SM 12.1)
@Sggin13mo ago
AI & ML
The official repo for the paper "Optimizing LLM Inference Throughput via Memory-aware and SLA-constrained Dynamic Batching"
@KevinLee11101y ago
AI & ML
Collection of recent advanced RAG techniques.
@hienhayho10mo ago
AI & ML
Source code for the paper: Hear Both Sides: Efficient Multi-Agent Debate via Diversity-Aware Message Retention
@DA2I2-SLM5mo ago
AI & ML
An OpenAI Compatible API which integrates LLM, Embedding and Reranker. LLM、Embedding Reranker OpenAI API
@hcd2331y ago
AI & ML
Drop-in OIDC & Google A2A auth + Weaviate memory for Ollama, vLLM and any local LLM server.
@attach-dev8mo ago
AI & ML
An open-source, model-agnostic agent harness for local LLMs. Define agents in YAML (tools, memory, deny-first permissions) and run them against any OpenAI-compatible endpoint: vLLM, Ollama, LM Studio, or llama.cpp.
@ahwurm4w ago
AI & ML
Agentic SDLC with local LLM's.
@xencon4w ago
AI & ML
Dual-mode cloud burst LLM router. Edge-first or cloud-first inference with data sovereignty, cost controls, and automatic failover. https://aiburstcloud.com
@aiburstcloud5w ago
AI & ML
Convert pdf and image files into markdown
@arc533mo ago
DevTools
htop for your LLM inference cluster
@InfraWhisperer4mo ago
AI & ML
KONASH: Train knowledge agents that search, retrieve, and reason. Based on KARL (Databricks, 2026).
@konaequity5mo ago
AI & ML
Open recipes, engine patches, and benchmark harnesses for LLM inference on Intel Arc Pro B60/B70 (Battlemage, Xe2). MoE 35B at 126 t/s decode / 7.5K t/s prefill single-stream — vLLM XPU MTP unlocked.
@SergiioB3w ago
AI & ML
Interested in running a conspiracy of agents (technical term) on your Local AI Infra? Who isn't!
@digitalspaceport5mo ago
AI & ML
vLLM + Qwen3.6-27B (BF16) OpenAI-compatible inference server on AMD Strix Halo (Ryzen AI Max+ 395, gfx1151). Vision input, 256K context, /v1/responses with separated reasoning, via TheRock ROCm.
@hec-ovi4mo ago
AI & ML
Diagnose vLLM inference servers
@vllm-doctor5w ago
DevTools
The official implementation of NeurlPS 2025 D&B paper: IndustryEQA: Pushing the frontiers of Embodied Question Answering in Industrial Scenarios.
@JackYFL11mo ago
DevTools
A project worth exploring.
@ampir-nn8mo ago
AI & ML
Hikari's curated English measured notes on local LLM systems (serving, specdec, training, steering).
@hikarioyama5w ago
AI & ML
Thinking-native compiler for LEET Language Understanding item sets — killer-item design grammar, verification separation, deterministic gate chain, 3-model benchmark
@Glockevonpavlov5w ago
AI & ML
A curated collection of Claude Code agent skills that accelerate the entire vLLM development lifecycle.
@shen-shanshan3w ago
AI & ML
TokenPowerBench: Benchmarking the power consumption of LLM inference across engines and deployment configurations
@chenxuniu5mo ago
AI & ML
Exocomp Agentic Environment for Go
@cookiengineer3mo ago
DevTools
Reproducible recipe: serve abliterated Gemma-4-12B (gemma4_unified) at 50-118 tok/s on no-NVLink Blackwell (SM120) via vLLM nightly + ModelOpt FP8/NVFP4 + MTP spec-decode.
@lna-lab3mo ago
AI & ML
AI-Inference-Managed-by-AI: Go binary for managing AI inference on edge devices
@Approaching-AI3w ago
DevTools
Hands-on GPU/HPC infrastructure operations: K8s GPU scheduling, HAMi sharing, Slurm, observability & vLLM inference. Learn it free on a laptop; validate on one cheap GPU.
@ld-singh5w ago
AI & ML
Easily boost the speed of pulling your models and datasets from various of inference runtimes. (e.g. HuggingFace, Ollama, vLLM, and more!)
@moeru-ai1y ago
AI & ML
FlashAttention-style custom attention backend for vLLM on AMD MI50/MI60/Radeon VII (gfx906). Downstream fork of mixa3607/ML-gfx906 with replacement HIP kernels and a vllm.general_plugins entry point.
@nick413-bit5mo ago
DevTools
Running Large Language Model easily.
@janelu95w ago
AI & ML
Turn any NVIDIA GPU into a local AI platform. Inference + fine-tuning in your browser. One command to start, automatic clustering.
@getainode2mo ago
AI & ML
Source-code level analysis of LLM RL training infra: async RL, weight sync, FP8, MoE routing | LLM RL
@zpqiu6w ago
AI & ML
FlashHead: Efficient Drop-In Replacement for the Classification Head in Language Model Inference
@embedl8w ago
AI & ML
SAM — Smart Agentic Model: CLI coding agent for open-source LLMs. pip install sam-agent
@SecFathy6mo ago
DevTools
Automated Triton w8a8 block FP8 kernel tuning tool for vLLM. Auto-detects model architecture, supports Qwen3-Coder-30B-A3B-Instruct-FP8/DeepSeek-V3/custom models, multi-GPU parallel tuning, and generates optimized kernel configs for quantization.
@massif-014mo ago
AI & ML
AI hardware fit calculator for LLM, embedding, reranker, OCR and VLM workloads — VRAM, throughput, licensing and multi-GPU planning
@jaeseok6143w ago
AI & ML
Local LLM speed measurements from a 4x RTX 3090 rig, with exact launch commands and per-run provenance.
@alesha-pro4w ago
SaaS
Multi-tenant streaming ASR server for Qwen3-ASR (vLLM AsyncLLMEngine). Drop-in replacement for qwen-asr-demo-streaming.
@jayter-official3mo ago
AI & ML
Honeycomb Lab — hex map + OpenAI gateway control plane for a home AI fleet
@joeynyc4w ago
AI & ML
LLM Performance Testing | K6 + Grafana + InfluxDB | A tiny toolkit for load testing and benchmarking OpenAI-like inference endpoints using K6 + Grafana + InfluxDB
@0xnyn1y ago
AI & ML
A high-throughput LLM serving engine with non-uniform KV cache compression, built on vLLM
@aiha-lab4w ago
AI & ML
A lean, local AI research agent framework that runs on your machine
@KabakaWilliam5mo ago
AI & ML
LLM, Fine Tuning, Llama 2, Gemma, Mixtral, vLLM, LangChain, RAG, ChromaDB, FAISS
@joydeb282y ago
AI & ML
Kubernetes scanner that discovers LLMs running on vLLM and extracts their deployment and runtime facts.
@paralleliq2mo ago
AI & ML
GPU Service Manager for LLM workloads on Linux/NVIDIA systems.
@jaigouk2mo ago
AI & ML
Local-first AI coding agent runtime — your models, fail-closed safety, evidence-gated self-improvement
@cdnwetzel3w ago
AI & ML
Batch LLM Inference with Ray Data LLM: From Simple to Advanced
@0-mostafa-rezaee-07mo ago
AI & ML
Benchmark OpenAI-compatible AI endpoints and AI Accelerators in a reproducible structured way
@flexaihq6mo ago
DevTools
vLLM tool parser for Qwen2.5-Coder models using <tools> tag format.
@hanXen4mo ago
AI & ML
Distributed Inference with vLLM
@KempnerInstitute7w ago
AI & ML
LLM inference engine built from scratch in C++. No PyTorch, no frameworks.
@Anirudh1712025mo ago
AI & ML
The Continuous Verification and Optimization layer for self-hosted LLMs
@colomalabs7w ago
AI & ML
An autonomous agentic pipeline that finds, proves, and patches real C memory-safety vulnerabilities end-to-end using a single 7B open-source LLM (Qwen2.5-Coder) on vLLM, with AddressSanitizer as a ground-truth oracle.
@FareedKhan-dev3mo ago
DevTools
GLM-5.2 744B with a true 4-bit NVFP4 KV cache on 4x DGX Spark: 42 tok/s peak, 317K-token KV pool (+58.6% vs fp8), needle-verified at 250K depth, serving 316K context
@tonyd2wild6w ago
DevTools
Open-source local AI server configs, GFX906 runtime maintenance, reproducible benchmarks, and QC methods for affordable AI research infrastructure.
@joe2gaan2mo ago
AI & ML
Ask Poddy: Run Open Source LLMs and Embeddings as OpenAI-Compatible Serverless Endpoints (Tutorial)
@blib-la2y ago
DevTools
An out-of-tree vLLM plugin for Mobilint NPU runtime integration.
@mobilint5w ago
AI & ML
Kubernetes-native control plane for scale-to-zero serving of long-tail LLMs
@noctaya4w ago
DevTools
GPU orchestrator with vLLM sleep mode
@arkorlab4w ago
AI & ML
Run DeepSeek-V4-Flash 304B on a single NVIDIA GB10 / DGX Spark - 2-bit MoE planes, FP4 quality recovery, speculative decoding
@lrozewicz3w ago
AI & ML
nano-vllmmoeSpeculative Decoding
@banfeb3mo ago
AI & ML
MinerU AMD vLLM + hybrid-auto-engine ! AMD/ROCm 。 MinerU 3.x ROCm 7.x + PyTorch 2.11.0 + vLLM , NVIDIA 。
@buptanswer3mo ago
DevTools
GLM-5.2 on 4x DGX Spark with adaptive MTP K2/K4/K5, FULL CUDA graphs, DCP2, 520K context, and a downloadable ARM64 runtime image.
@0xdfi4w ago
AI & ML
Tencent Hy3 295B MoE (NVFP4) on 2x NVIDIA DGX Spark — TP2 over 200GbE, 256K context, 26 tok/s end-to-end. Tuned, benchmarked, agent-ready.
@joeynyc2mo ago
DevTools
Professional CLI toolkit for Modal GPU workflows, Docker Compose deployments, Python runtimes, model serving, and account/billing management.
@PuxHocDL6w ago
AI & ML
Qwen3.6 vLLM Toolkit — Launcher + Templates, Optimized for 2×24 GB. vLLM launcher with an optimized hybrid chat template drawing from the best community fixes for Qwen 3.6 27B.
@pajitosingh7w ago
AI & ML
Terraform setup for deploying a private coding LLM on Vast.ai with vLLM, Qwen3 Coder, and OpenCode.
@anqorithm4mo ago
AI & ML
Extend LLM context windows beyond GPU memory limits with disk-backed KV cache.
@tirdyhouse3mo ago
DevTools
npm pacakge to Craft files into Markdown with ease
@jparkerweb1y ago
DevTools
General Information, model certifications, and benchmarks for nm-vllm enterprise distributions
@neuralmagic1y ago
AI & ML
Multi-model LLM serving for NVIDIA DGX Spark with vLLM, web UI, and tool calling
@dataforgex7mo ago
DevTools
The official implementation of CVPR Workshop 2025 paper: Window Token Concatenation for Efficient Visual Large Language Models.
@JackYFL1y ago
AI & ML
agentsculptor is an experimental AI-powered development agent designed to analyze, refactor, and extend Python projects automatically. It uses an OpenAI-like planner–executor loop on top of a vLLM backend, combining project context analysis, structured tool calls, and iterative refinement. It has only been tested with gpt-oss-120b via vLLM.
@Perpetue23712mo ago
AI & ML
A slim, local-LLM-first AI coding assistant. TUI, Web UI, and desktop app. Runs air-gapped with Ollama, vLLM, or any OpenAI-compatible endpoint.
@bobbyjohnstx4w ago
DevTools
Spark Pulse is a web control plane for spark-vllm-docker
@kharkevich-engineering-lab3mo ago
AI & ML
Tiny, stateless Go router that dispatches OpenAI-compatible requests to single-model vLLM and sglang backends with zero external dependencies
@mirkolenz4w ago
AI & ML
Repository for Multililngual Generation, RAG evaluations, and surrogate judge training for Arena RAG leaderboard (NAACL'25)
@vectara1y ago
DevTools
Docker Compose setup that integrates vLLM, Open WebUI and NGINX (for SSL termination).
@marib001y ago
AI & ML
A voice assistant with local LLM as a backend
@Saga91031y ago
AI & ML
Open-source, self-hosted LLM chat application, featuring local-first data storage and real-time streaming responses.
@GHuyHuynh11mo ago
AI & ML
A curated collection of NLP and LLM resources. Covers essential papers and blogs on Transformers, Reinforcement Learning (RLHF, DPO, GRPO), Mechanistic Interpretability, Scaling Laws, and MLSys.
@rraghavkaushik4mo ago
AI & ML
Automated comic cataloging tool that identifies issues directly from cover images using a vision-language model, then cross-references results with the Grand Comics Database and the ComicVine API to generate structured, high-confidence collection data with minimal manual entry.
@boyobob6mo ago
AI & ML
AI-,. XTTS v2, real-time (Vosk/Whisper) offline LLM. - (Vue 3), Telegram-,, fine-tuning pipeline. Self-hosted,,.
@ShaerWare3w ago
AI & ML
VS Code extension to connect to local LLM inference servers (Ollama, llama.cpp, LocalAI) via OpenAI API compatible endpoints.
@krevas4w ago
AI & ML
Read PGS (Bluray) and VobSub (DVD) image subtitles and extract their text using external Vision Language Models.
@hekmon6w ago
AI & ML
Perl Framework for AI - Langertha - the viking of AI
@Getty3w ago
DevTools
Lna-Lab production pipeline: GGUF -> modelopt-format NVFP4 + working MTP head for vLLM on RTX PRO 6000 Blackwell (SM120). Stages 2 (NVFP4) and 3 (MTP graft) are Lna-Lab originals; stage 1 (GGUF->bf16) reuses li-yifei/gguf-to-nvfp4.
@lna-lab4mo ago
AI & ML
Benchmarks and notes for running modern LLMs with vLLM on 8x Tesla V100-32GB in 2026.
@KumphanartDansiri8w ago
AI & ML
inference cookbook / inference
@SuperMarioYL6mo ago
AI & ML
GLM-5.2 QuantTrio (Unpruned ) on 4xDGX Spark (GB10): 327K-655K context or up to 5 concurrent agents, one receipe. TP4 + DCP + MTP
@XanuNetworks7w ago
DevTools
nvtop for vLLM — an interactive terminal dashboard for vLLM serving performance (concurrency, throughput, cache & KV memory, latency, spec-decode, GPU)
@bryanvine2mo ago

Ask MadeWithWhat

AI answers may contain mistakes — please double-check important details.