vLLM
PaketOpen-source engine for serving LLMs with high throughput on GPUs
- Fiyat
- Free, open source
- Erişim
- None; optional server API key
Hakkında
Python inference engine that serves open-weight models behind an OpenAI-compatible server, using PagedAttention and continuous batching. Built for teams hosting models on their own GPUs. Linux is the main platform, with NVIDIA, AMD, TPU and other backends.
Neler yapabilirsin
- Serve a Hugging Face model as an OpenAI-compatible API on your own GPUs
- Run offline batch inference over large prompt sets from Python
- Scale serving across GPUs and nodes with tensor and pipeline parallelism
Başlarken
- Install on Linux: uv pip install vllm --torch-backend=auto
- Start the server with vllm serve and a Hugging Face model id
Örnek kod
vllm serve Qwen/Qwen2.5-1.5B-Instruct
# in another terminal
curl http://localhost:8000/v1/chat/completions \
-H "Content-Type: application/json" \
-d '{
"model": "Qwen/Qwen2.5-1.5B-Instruct",
"messages": [{"role": "user", "content": "Who won the world series in 2020?"}]
}'Ayrıntılar
- Barındırma
- Kendi sunucunda, Kendi bilgisayarında
- Kullanılabildiği yerler
- Tüm dünya
- Resmi SDK'lar
- Python
- MCP sunucusu
- Yok