< cd ../gallery
[AI & ML]
dynamic-batching
The official repo for the paper "Optimizing LLM Inference Throughput via Memory-aware and SLA-constrained Dynamic Batching"
@KevinLee1110
maintainer
★ 18 stars
# README
The official repo for the paper "Optimizing LLM Inference Throughput via Memory-aware and SLA-constrained Dynamic Batching". Maintained by KevinLee1110 on GitHub, where it has earned 18 stars from the community.
It's actively developed around inference-serving, llm, vllm, and is a solid reference for anyone building with these tools.
# tags
# install
npm install dynamic-batchingmore vLLM repos