Ink: The fastest and most accurate speech to text model
Join the teams making the switch to Cartesia
One transcription model for every environment your business takes you to
Trains rumble past, announcements crackle overhead. Ink-2 transcribes every word the caller says.
Noisy city
Built for Voice Agents
Four capabilities that make Ink the transcription layer production agents rely on.
Accuracy
Heard right the first time.
In practice
In a voice agent, the transcript is the foundation everything else builds on. A transcription error undermines the LLM input and takes the interaction in the wrong direction.
The inverse is equally true — accuracy compounds, and a precise transcript means a better response and a call that resolves.
Ink-2's approach
Ink has the lowest Word Error Rate (WER) of any streaming STT model, natively handling structured data — phone numbers, dates, emails, currencies, and UUIDs. Built for real-world audio settings — telephony, background noise, varied accents, and more.
Dates, alphanumerics, IDs
Conversational flow
Knows when you start and finish.
In practice
A conversation has two critical moments — when a caller starts talking and when they finish. Miss the start and the agent misses the turn entirely. Trigger too early on the end and the agent jumps in mid-thought. The right transcription model gets both right without the wait.
Ink-2's approach
Ink-2 is built with native turn detection — turn.start and turn.end signaled directly by the model, with no external VAD to integrate or maintain. For lower latency, turn.eager_end gives your LLM a head start before the turn is confirmed complete.
Semantic endpointing determines turn end by meaning, not silence — so pauses mid-thought don't trigger the agent prematurely.
ink-2
Your agent can respond before the competitor even starts
flux-general-en
Only now does the agent know the caller finished — it starts from scratch
Speed
The caller stops talking.
The agent starts thinking.
In practice
When transcription is fast and consistent, the agent's response feels immediate. One slow transcript in ten means one call in ten where that readiness breaks. Nine great calls don't cancel out the one that didn't feel right.
Ink-2's approach
Ink is the fastest streaming ASR model - built on a custom inference engine purpose-built for real-time conversation. Time to final transcript is 0.1s, with turn.eager_end reducing the gap between the last word and the first response.
Cost
Quality that doesn't cost more as you grow.
In practice
Voice is the most natural interface for communication. Getting cost and quality right at scale enables voice everywhere — the default interface across every agentic interaction.
Ink-2's approach
Ink's State Space Model architecture delivers 10-100x the throughput of transformers — lower compute cost at scale, with no quality tradeoffs. Ongoing optimization of our model stack means better unit economics as you scale.
Enterprise-grade security.From Cloud to Local.
HIPAA compliant

SOC 2 Type 2

GDPR

PCI
FAQs
Why is Ink-2 the best STT for voice agents?
Ink-2 is Cartesia's streaming speech-to-text model, purpose-built for production voice agents. Like our TTS, it's built on State Space Model (SSM) architecture pioneered by our founding team at Stanford, which is what enables Ink-2 to deliver three things at once:
Lowest WER of any streaming STT model. Ink-2 outperforms Deepgram Flux, Soniox RT-V4, AssemblyAI RT Pro, ElevenLabs Scribe-2-realtime, and other production streaming models across line recordings, accented speech, noisy conditions, and earnings calls, including alphanumerics like phone numbers, emails, and UUIDs.
Best-in-class turn detection, built in. Ink-2 has model-integrated end-of-turn detection, meaning you don't need a separate turn-taking model like Silero in your stack. Fewer dependencies, lower latency, more accurate turn-taking.
Built-in noise robustness. Ink-2 is robust to background noise without requiring Krisp or other audio filters, removing cost and latency from your stack.
Do I still need Krisp or Silero with Ink-2?
No. This is one of the main reasons teams switch to Ink-2:
Turn-taking is built in. Most STT models require an external turn detector (like Silero) to know when the user is done speaking. Ink-2 handles this natively, which is both faster and more accurate than running a separate model.
Noise robustness is built in. Most voice agents stack Krisp or similar audio filters on top of their STT to remove background noise, adding cost, latency, and complexity. Ink-2 is robust to background noise out of the box.
Fewer models in your stack means lower latency, lower cost, and fewer points of failure.
Can Ink-2 run on-prem or in my own cloud (VPC)?
Yes, Ink-2 can be deployed:
On-prem inside your data center, including air-gapped environments
In your own VPC on AWS, GCP, or Azure
Via OEM licensing for embedding Ink-2 directly into your product
This makes Ink-2 viable for government contracting, regulated industries (healthcare, financial services, insurance), and customers with data sovereignty or residency requirements. On-prem and OEM deployments are available under enterprise contracts.
What languages does Ink-2 support?
Ink-2 supports English, Spanish, French, Hindi, and Japanese.
How much does Ink-2 cost?
Ink-2 is priced at 3 credits/second. Pricing details and volume discounts are on the pricing page at https://www.cartesia.ai/pricing. Enterprise contracts include custom pricing based on usage and deployment model.
Cartesia also offers a startup grants program at https://www.cartesia.ai/startups with credits for qualifying early-stage companies.
When should I contact Sales?
Reach out to the Cartesia team if any of the following apply:
You're running high-volume production workloads (>50M credits)
You need on-prem, VPC, or OEM deployment
You need a BAA, zero data retention, or other contractual compliance terms for healthcare, financial services, or regulated sectors
You're in government, federal, or public sector procurement
For everything else — evaluation, prototyping, smaller production workloads — the self-serve plans & technical documentation at https://docs.cartesia.ai/build-with-cartesia/stt/latest will get you started.
Get started today
Talk to an expert.
Connect with a member of our team and learn how Cartesia can help you build world-class voice experiences.
Start building.
Access our models via API and bring a voice agent into production in minutes.





