< cd ../gallery
[DevTools]

vllm_benchmark_block_fp8

Automated Triton w8a8 block FP8 kernel tuning tool for vLLM. Auto-detects model architecture, supports Qwen3-Coder-30B-A3B-Instruct-FP8/DeepSeek-V3/custom models, multi-GPU parallel tuning, and generates optimized kernel configs for quantization.
massif-01
@massif-01
maintainer
★ 13 stars
go to website ↗github
vllm-benchmark-block-fp8 — preview
vllm_benchmark_block_fp8 repository preview
# README
Automated Triton w8a8 block FP8 kernel tuning tool for vLLM. Auto-detects model architecture, supports Qwen3-Coder-30B-A3B-Instruct-FP8/DeepSeek-V3/custom models, multi-GPU parallel tuning, and generates optimized kernel configs for quantization. Maintained by massif-01 on GitHub, where it has earned 13 stars from the community.
It's actively developed around fp8, kernel-tuning, performance-tuning, and is a solid reference for anyone building with these tools.
# tags
# install
npm install vllm_benchmark_block_fp8
languages
Python100%
last commit2 months ago
license
more vLLM repos
$ made-with-vllm
rssllmmadewithwhat

Ask MadeWithWhat

AI answers may contain mistakes — please double-check important details.