Cartesia

Model

Real-time voice models for text to speech and speech to text

Price
Free tier, then paid
Access
API key

About

Voice AI lab with Sonic text to speech, Ink speech to text and hosted voice agents, built for low-latency streaming. Plans include monthly credits, and speech costs about 1 credit per character. Ink 2 supports five languages.

What you can do with it

  • Stream speech from LLM output over a WebSocket with low latency
  • Transcribe live audio with built-in turn detection
  • Build and deploy voice agents with telephony and tool calls

Get started

  1. Create an API key at play.cartesia.ai/keys
  2. Export it as CARTESIA_API_KEY

Example

curl -X POST https://api.cartesia.ai/tts/bytes \
  -H "Authorization: Bearer $CARTESIA_API_KEY" \
  -H "Cartesia-Version: 2026-08-14" \
  -H "Content-Type: application/json" \
  -d '{"model_id":"sonic-3.6","transcript":"Hi there! Welcome to Cartesia Sonic.","voice":"f786b574-daa5-4673-aa0c-cbe3e8534c02","output_format":{"container":"wav","encoding":"pcm_s16le","sample_rate":44100}}' \
  --output sonic.wav

Details

Hosting
Hosted service
Available in
Worldwide
Official SDKs
JavaScript/TypeScript, Python
MCP server
Remote

Tasks

Alternatives

Other tools for the same tasks.

Last checked on .