Hugging Face Inference
ProviderOne API for open models served by many inference providers
- Price
- Free tier, then paid
- Access
- User access token
About
Hugging Face's serverless API that routes requests for open-weight models to partner providers such as Groq, Together and fal, with one token and one bill. Suits developers who want many open models without separate accounts. The monthly free credit is small.
What you can do with it
- Call open-weight LLMs like gpt-oss and DeepSeek through one OpenAI-compatible endpoint
- Pick the fastest or cheapest provider per model with a suffix like :cheapest
- Generate images, embeddings and transcripts from Hub models in Python or JS
Get started
- Create a fine-grained token with Inference Providers permission
- Pick a model that a provider serves on the Hub
Example
curl https://router.huggingface.co/v1/chat/completions \
-H "Authorization: Bearer $HF_TOKEN" \
-H "Content-Type: application/json" \
-d '{
"model": "openai/gpt-oss-120b:fastest",
"messages": [{"role": "user", "content": "How many G in huggingface?"}],
"stream": false
}'Details
- Hosting
- Hosted service
- Available in
- Worldwide
- Official SDKs
- Python, JavaScript/TypeScript
- MCP server
- None