vLLM

Package

Open-source engine for serving LLMs with high throughput on GPUs

Price
Free, open source
Access
None; optional server API key

About

Python inference engine that serves open-weight models behind an OpenAI-compatible server, using PagedAttention and continuous batching. Built for teams hosting models on their own GPUs. Linux is the main platform, with NVIDIA, AMD, TPU and other backends.

What you can do with it

  • Serve a Hugging Face model as an OpenAI-compatible API on your own GPUs
  • Run offline batch inference over large prompt sets from Python
  • Scale serving across GPUs and nodes with tensor and pipeline parallelism

Get started

  1. Install on Linux: uv pip install vllm --torch-backend=auto
  2. Start the server with vllm serve and a Hugging Face model id

Example

vllm serve Qwen/Qwen2.5-1.5B-Instruct
# in another terminal
curl http://localhost:8000/v1/chat/completions \
  -H "Content-Type: application/json" \
  -d '{
    "model": "Qwen/Qwen2.5-1.5B-Instruct",
    "messages": [{"role": "user", "content": "Who won the world series in 2020?"}]
  }'

Details

Hosting
Self-hosted, Runs locally
Available in
Worldwide
Official SDKs
Python
MCP server
None

Tasks

Alternatives

Other tools for the same tasks.

Last checked on .