Cartesia

Research Team Data Published Compare

Summary

Real-time speech and transcription model platform for voice agents, offering Sonic TTS, Ink STT, and enterprise voice-agent APIs.

Description

Cartesia provides Sonic text-to-speech, Ink speech-to-text, voice-agent APIs and real-time speech models optimized for low latency, quality, enterprise-grade scale and integration into live voice-agent systems.

Positioning

Real-time speech models purpose-built for voice agents

Key facts

HQ location
San Francisco, CA, USA
Founded
2023
Employee range
51-200
Funding stage
Series A
Company type
Private
Pricing model
Freemium Usage Based (Free tier; Pro $5/mo; usage-based credits/agent minutes; enterprise plans)
Last updated

Financials

Revenue estimate
Unknown / not publicly disclosed
Valuation estimate
Unknown
Investments
$64M Series A announced Mar 2025; $91M total funding reported; later 2025/2026 round reports unverified

Relationships

Target customers
Developers, AI agent companies and enterprises building real-time voice agents
Key competitors
ElevenLabs, Deepgram, Hume AI, OpenAI, PlayHT
Known customers
Quora, Cresta, Rasa, ServiceNow, Sierra

Classification (raw research text)

Core focus
Voice AI infrastructure
Core industry
AI infrastructure / voice AI
Core category
STT/TTS and voice agent infrastructure

Shown verbatim from the research spreadsheet — deriving structured industry tags from this text is a future phase.

Segments, Industries & Certifications

Segments

AI Workflows, AI Agents, LLM Fine-tuning, AI Developer Tools, Knowledge & RAG, LLM Deployment, AI Quality & Observability, AI Governance & Risk, Call Analysis, Voice AI Agents, STT/TTS Infrastructure, Analytics & BI