Speech to text, live in your browser

Click the mic, start talking, and watch the transcript appear. It runs on Ink-2, ranked #1 for streaming speech-to-text accuracy.
Click or press Space to start a transcription
Tap to start a transcription
Need a topic?
What's a place you'd love to visit someday?

Powered by Ink-2, ranked #1 for streaming accuracy

The demo above runs on Ink-2, our streaming speech-to-text model. It ranks first on the VoiceArena US-English leaderboard and Artificial Analysis' streaming benchmark, and it transcribes structured data like phone numbers, emails, and UUIDs right the first time. Turn detection is built in, so the transcript breaks into turns on its own, and it stays accurate in background noise without a separate cleanup filter.

Good for

  • Voice agents
  • Live captions
  • Dictation
  • Meeting notes
  • Call analytics

How it works

01

Click the mic and talk. Words appear as you speak, no sign-up required.

02

Pause or switch topics mid-stream. Ink-2 marks where turns start and end, so the transcript reads like a conversation.

03

Copy the transcript, or stream the same model from your app through the API.

Take it to production

The model behind this page is the same Ink-2 you get from our API and SDKs. Stream live audio over a WebSocket for real-time use, or send finished files to the batch endpoint. Keyterm prompting biases recognition toward the names and jargon your product hears most. Enterprise plans add on-prem, VPC, and OEM deployments with zero data retention.

FAQs

What is speech to text?

Speech to text (also called STT, ASR, or transcription) turns spoken audio into written words. Cartesia's version runs on Ink-2, a streaming model, so the transcript appears while you are still talking instead of after you stop.

Is this speech to text tool free?

Yes. The demo on this page transcribes your microphone without an account. A free Cartesia account raises the limits, and paid plans add production API volume.

What languages does it support?

Ink-2 transcribes English, French, Hindi, Japanese, and Spanish, and detects the language automatically. It is built for real-time use; for pre-recorded files, the batch STT API runs on ink-whisper.

What is turn detection?

Turn detection decides when a speaker is done, which is what keeps a voice agent from interrupting or dead-airing. Most stacks bolt on a separate voice-activity model for this. Ink-2 emits turn events like turn.eager_end and turn.end itself, which means one fewer model in the stack and faster responses.

Can I transcribe audio files instead of my microphone?

The demo on this page is live and microphone-driven. For recordings, use the batch STT API, which accepts a complete file and returns one transcript.

Can I run it in my own infrastructure?

Yes. Ink-2 deploys on-prem (including air-gapped environments), in your own VPC on AWS, GCP, or Azure, or as an OEM license embedded in your product. These deployments are available under enterprise contracts.

Get started today

Talk to an expert.

Connect with a member of our team and learn how Cartesia can help you build world-class voice experiences.

Contact Sales

Start building.

Access our models via API and bring a voice agent into production in minutes.

Try Cartesia