vLLM
PackageOpen-source engine for serving LLMs with high throughput on GPUs
- Price
- Free, open source
- Access
- None; optional server API key
About
Python inference engine that serves open-weight models behind an OpenAI-compatible server, using PagedAttention and continuous batching. Built for teams hosting models on their own GPUs. Linux is the main platform, with NVIDIA, AMD, TPU and other backends.
What you can do with it
- Serve a Hugging Face model as an OpenAI-compatible API on your own GPUs
- Run offline batch inference over large prompt sets from Python
- Scale serving across GPUs and nodes with tensor and pipeline parallelism
Get started
- Install on Linux: uv pip install vllm --torch-backend=auto
- Start the server with vllm serve and a Hugging Face model id
Example
vllm serve Qwen/Qwen2.5-1.5B-Instruct
# in another terminal
curl http://localhost:8000/v1/chat/completions \
-H "Content-Type: application/json" \
-d '{
"model": "Qwen/Qwen2.5-1.5B-Instruct",
"messages": [{"role": "user", "content": "Who won the world series in 2020?"}]
}'Details
- Hosting
- Self-hosted, Runs locally
- Available in
- Worldwide
- Official SDKs
- Python
- MCP server
- None