Cartesia vs Fish Audio
Fish Audio's numbers are fishy.
Fish Audio claims a 67% human preference rate over competitors and that they're twice as fast as Cartesia. Every independent benchmark says otherwise.


Quality Elo, voice-matched
Higher is better
Voice quality
Independent benchmarks rate Sonic 3.5 higher
- Fish claims a 67% blind-test preference over leading competitors, and says listeners can't reliably distinguish S2.1 Pro from human across 581 published head-to-head comparisons.
- Independent benchmarks say otherwise.
Latency, P90
Lower is better
Latency
Cartesia is 2x faster than Fish on P90 latency
- Marketing quotes the fastest request. Users experience tails.
- Cartesia's P90 is 394ms. Fish's is 727ms. 2x faster, and not in the direction Fish claims.
Reasons why companies choose Cartesia over Fish Audio
Human-like naturalness
Intonation, pacing, pronunciation, emotion, and audio quality that drive higher completion rates
Cartesia1104 EloAugust 2026
Fish Audio1005 EloAugust 2026
Latency (P90)
Decides whether it feels like a conversation or a phone tree.
Cartesia394msJuly 2026
Fish Audio727msJuly 2026
Cost
Great models don't need to be expensive
Cartesia~50% cheaper than ElevenLabs at higher quality, SSM architecture is more compute & cost efficientFish AudioCheaper on paper, steep quality drop on independent benchmarks, transformer-based worse compute efficiency at scaleEmotion
Moves escalation rates and satisfaction scores.
CartesiaModel-level emotion inference from context across all languages, optional SSML tags for explicit controlFish AudioWeak model-level inference, so emotion depends on the LLM picking the right tag, SSML tags only work on 13 languagesSelf-hosting
Decides whether regulated teams can use it at all.
CartesiaVPC, on-prem, on-device, or air-gapped available all on current flagship modelsFish AudioSelf-hosting rests on open weights of older models (S2, not S2.1 Pro), with no on-device path at allEnd-to-end stack
How many vendors it takes to run one call.
CartesiaTTS, STT and Agents. Only provider ranked #1 on both speech and transcriptionFish AudioTTS-first, with an STT that doesn't compete at the frontierPVCs
How much audio you have to collect before a custom voice is usable.
CartesiaRequires just 30 minutes of audioFish AudioRequires at least an hour of audioLanguages
Comprehension and trust across a global customer base.
Cartesia40+ languages, best-in-class quality on each one, available on all tiersFish Audio83 languages claimed, but no published list and their own docs tier only 3 as top quality
Frequently asked questions
Get started today
Talk to an expert. Connect with a member of our team and learn how Cartesia can help you build world-class voice experiences.
Contact SalesStart building. Access our models via API and bring an agent into production with our robust SDKs and developer tools.
Try CartesiaCapabilities