< cd ../gallery
[AI & ML]
imp
From-scratch C++/CUDA inference engine for the NVIDIA RTX 5090 (sm_120a) — the best single-GPU backend for agentic AI: tool calling, long-context loops, reasoning and concurrent sub-agents on top of the fastest single-stream decode on the 5090 (beats llama.cpp, at-or-ahead of vLLM on NVFP4). 100% written by Claude Code.
@kekzl
maintainer
★ 30 stars
# README
From-scratch C++/CUDA inference engine for the NVIDIA RTX 5090 (sm_120a) — the best single-GPU backend for agentic AI: tool calling, long-context loops, reasoning and concurrent sub-agents on top of the fastest single-stream decode on the 5090 (beats llama.cpp, at-or-ahead of vLLM on NVFP4). 100% written by Claude Code. Maintained by kekzl on GitHub, where it has earned 30 stars from the community.
It's actively developed around blackwell, cpp, cuda, and is a solid reference for anyone building with these tools.
# tags
# install
npm install impmore vLLM repos