Summary
Real-time speech and transcription model platform for voice agents, offering Sonic TTS, Ink STT, and enterprise voice-agent APIs.
Description
Cartesia provides Sonic text-to-speech, Ink speech-to-text, voice-agent APIs and real-time speech models optimized for low latency, quality, enterprise-grade scale and integration into live voice-agent systems.
Positioning
Real-time speech models purpose-built for voice agents
Key facts
- HQ location
- San Francisco, CA, USA
- Founded
- 2023
- Employee range
- 51-200
- Funding stage
- Series A
- Company type
- Private
- Pricing model
- Freemium Usage Based (Free tier; Pro $5/mo; usage-based credits/agent minutes; enterprise plans)
- Last updated
Financials
- Revenue estimate
- Unknown / not publicly disclosed
- Valuation estimate
- Unknown
- Investments
- $64M Series A announced Mar 2025; $91M total funding reported; later 2025/2026 round reports unverified
Relationships
- Target customers
- Developers, AI agent companies and enterprises building real-time voice agents
- Key competitors
- ElevenLabs, Deepgram, Hume AI, OpenAI, PlayHT
- Known customers
- Quora, Cresta, Rasa, ServiceNow, Sierra
Classification (raw research text)
- Core focus
- Voice AI infrastructure
- Core industry
- AI infrastructure / voice AI
- Core category
- STT/TTS and voice agent infrastructure
Shown verbatim from the research spreadsheet — deriving structured industry tags from this text is a future phase.
Segments, Industries & Certifications
- Segments