Hugging Face Inference

Provider

One API for open models served by many inference providers

Price
Free tier, then paid
Access
User access token

About

Hugging Face's serverless API that routes requests for open-weight models to partner providers such as Groq, Together and fal, with one token and one bill. Suits developers who want many open models without separate accounts. The monthly free credit is small.

What you can do with it

  • Call open-weight LLMs like gpt-oss and DeepSeek through one OpenAI-compatible endpoint
  • Pick the fastest or cheapest provider per model with a suffix like :cheapest
  • Generate images, embeddings and transcripts from Hub models in Python or JS

Get started

  1. Create a fine-grained token with Inference Providers permission
  2. Pick a model that a provider serves on the Hub

Example

curl https://router.huggingface.co/v1/chat/completions \
  -H "Authorization: Bearer $HF_TOKEN" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "openai/gpt-oss-120b:fastest",
    "messages": [{"role": "user", "content": "How many G in huggingface?"}],
    "stream": false
  }'

Details

Hosting
Hosted service
Available in
Worldwide
Official SDKs
Python, JavaScript/TypeScript
MCP server
None

Tasks

Alternatives

Other tools for the same tasks.

Last checked on .