Groq
ProviderLow-latency inference API for open LLMs and Whisper
- Price
- Free tier, then paid
- Access
- API key
About
GroqCloud runs open models such as GPT-OSS, Qwen and Whisper on Groq's inference cloud behind an OpenAI-compatible API. It is for apps that need very fast token output. Its model list is short compared with larger hosts.
What you can do with it
- Run GPT-OSS and Qwen models with very low latency
- Transcribe audio fast with Whisper Large v3 Turbo
- Switch from OpenAI by changing the base URL to Groq
Get started
- Create an API key in the GroqCloud console
- Set GROQ_API_KEY in your environment
Example
curl -X POST https://api.groq.com/openai/v1/responses \
-H "Authorization: Bearer $GROQ_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "openai/gpt-oss-20b",
"input": "Explain the importance of fast language models"
}'Details
- Hosting
- Hosted service
- Available in
- Worldwide
- Official SDKs
- Python, JavaScript/TypeScript
- MCP server
- None