< cd ../gallery
[AI & ML]

imp

From-scratch C++/CUDA inference engine for the NVIDIA RTX 5090 (sm_120a) — the best single-GPU backend for agentic AI: tool calling, long-context loops, reasoning and concurrent sub-agents on top of the fastest single-stream decode on the 5090 (beats llama.cpp, at-or-ahead of vLLM on NVFP4). 100% written by Claude Code.
kekzl
@kekzl
maintainer
★ 30 stars
github
imp — preview
imp repository preview
# README
From-scratch C++/CUDA inference engine for the NVIDIA RTX 5090 (sm_120a) — the best single-GPU backend for agentic AI: tool calling, long-context loops, reasoning and concurrent sub-agents on top of the fastest single-stream decode on the 5090 (beats llama.cpp, at-or-ahead of vLLM on NVFP4). 100% written by Claude Code. Maintained by kekzl on GitHub, where it has earned 30 stars from the community.
It's actively developed around blackwell, cpp, cuda, and is a solid reference for anyone building with these tools.
# tags
# install
npm install imp
languages
Cuda54.972%
C++37.257%
Python4.684%
Shell2.355%
CMake0.368%
C0.159%
Makefile0.131%
Gnuplot0.002%
last commit1 week ago
licenseMIT
more vLLM repos
$ made-with-vllm
rssllmmadewithwhat

Ask MadeWithWhat

AI answers may contain mistakes — please double-check important details.