vLLM

Paket

Open-source engine for serving LLMs with high throughput on GPUs

Fiyat
Free, open source
Erişim
None; optional server API key

Hakkında

Python inference engine that serves open-weight models behind an OpenAI-compatible server, using PagedAttention and continuous batching. Built for teams hosting models on their own GPUs. Linux is the main platform, with NVIDIA, AMD, TPU and other backends.

Neler yapabilirsin

  • Serve a Hugging Face model as an OpenAI-compatible API on your own GPUs
  • Run offline batch inference over large prompt sets from Python
  • Scale serving across GPUs and nodes with tensor and pipeline parallelism

Başlarken

  1. Install on Linux: uv pip install vllm --torch-backend=auto
  2. Start the server with vllm serve and a Hugging Face model id

Örnek kod

vllm serve Qwen/Qwen2.5-1.5B-Instruct
# in another terminal
curl http://localhost:8000/v1/chat/completions \
  -H "Content-Type: application/json" \
  -d '{
    "model": "Qwen/Qwen2.5-1.5B-Instruct",
    "messages": [{"role": "user", "content": "Who won the world series in 2020?"}]
  }'

Ayrıntılar

Barındırma
Kendi sunucunda, Kendi bilgisayarında
Kullanılabildiği yerler
Tüm dünya
Resmi SDK'lar
Python
MCP sunucusu
Yok

Görevler

Alternatifler

Aynı görevler için başka araçlar.

Son kontrol: .