< cd ../gallery
[AI & ML]

vllm-awq4-qwen

vLLM Qwen 3.6-27B (AWQ-INT4) + DFlash speculative decoding on AMD Strix Halo (gfx1151 iGPU, 128 GB UMA, ROCm 7.13). 24.8 t/s single-stream, vision, tool calling, 256K context, OpenAI-compatible, Docker. Matches DGX Spark FP8+DFlash+MTP at a third of the cost. No CUDA.
hec-ovi
@hec-ovi
maintainer
★ 47 stars
github
vllm-awq4-qwen — preview
vllm-awq4-qwen repository preview
# README
vLLM Qwen 3.6-27B (AWQ-INT4) + DFlash speculative decoding on AMD Strix Halo (gfx1151 iGPU, 128 GB UMA, ROCm 7.13). 24.8 t/s single-stream, vision, tool calling, 256K context, OpenAI-compatible, Docker. Matches DGX Spark FP8+DFlash+MTP at a third of the cost. No CUDA. Maintained by hec-ovi on GitHub, where it has earned 47 stars from the community.
It's actively developed around 27b, amd-strix-halo, awq, and is a solid reference for anyone building with these tools.
# tags
# install
npm install vllm-awq4-qwen
languages
Python100%
last commit2 months ago
licenseUnlicense
more vLLM repos
$ made-with-vllm
rssllmmadewithwhat

Ask MadeWithWhat

AI answers may contain mistakes — please double-check important details.