Groq

Provider

Low-latency inference API for open LLMs and Whisper

Price
Free tier, then paid
Access
API key

About

GroqCloud runs open models such as GPT-OSS, Qwen and Whisper on Groq's inference cloud behind an OpenAI-compatible API. It is for apps that need very fast token output. Its model list is short compared with larger hosts.

What you can do with it

  • Run GPT-OSS and Qwen models with very low latency
  • Transcribe audio fast with Whisper Large v3 Turbo
  • Switch from OpenAI by changing the base URL to Groq

Get started

  1. Create an API key in the GroqCloud console
  2. Set GROQ_API_KEY in your environment

Example

curl -X POST https://api.groq.com/openai/v1/responses \
  -H "Authorization: Bearer $GROQ_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "openai/gpt-oss-20b",
    "input": "Explain the importance of fast language models"
  }'

Details

Hosting
Hosted service
Available in
Worldwide
Official SDKs
Python, JavaScript/TypeScript
MCP server
None

Tasks

Alternatives

Other tools for the same tasks.

Last checked on .