vllm-project/vllm

A high-throughput and memory-efficient inference and serving engine for LLMs

View on GitHub
Python
Stars 86.9k
Forks 19.7k
License Apache-2.0
Open Issues 5963
Updated 1h ago