Pipecat is an open-source Python framework for real-time voice agents, and Cartesia ships first-party speech services for it: Sonic for streaming text to speech and Ink for speech to text. In this guide you will run a working voice agent in your browser with both wired in, then use Ink’s eager turn-end predictions to shave roughly half a second off every response.
The whole build is a clone and two environment variables. The voice agent you end up with listens, decides when you have finished speaking, thinks, and answers out loud.
Overview
A Pipecat voice agent is a pipeline of processors. Audio enters through a transport, gets transcribed, passes through an LLM, and comes back out as speech:
transport.input() → STT → LLM → TTS → transport.output()
Two hops decide whether a conversation feels human or like a walkie-talkie: knowing when the user actually stopped talking, and speaking the reply fast. That is where the Cartesia services earn their place. CartesiaTurnsSTTService runs Ink 2, Cartesia’s streaming STT model, and lets the server decide when a turn is over instead of guessing locally. CartesiaTTSService streams Sonic audio as it generates, with word timestamps if you want them.
Pipecat’s own documentation puts a typical pipeline round trip at 500 to 800 milliseconds. The steps below are about spending that budget on the model rather than on silences.
Prerequisites
- Python 3.11 or later (the Pipecat repo’s README lists the current baseline).
- A Cartesia API key: create one in the dashboard.
- An LLM API key for the brain. The example uses OpenAI (
OPENAI_API_KEY); Pipecat also ships services for Anthropic, Gemini, Groq, Ollama, and others if you want a different model.
One scope note: Ink 2 is English-only today, so this agent listens in English. Sonic speaks more than 40 languages, so the reply side is not the constraint.
1. Install with the extras the example imports
The Pipecat repo has a ready-made Cartesia agent at examples/voice/voice-cartesia-turns.py:
git clone https://github.com/pipecat-ai/pipecat.git
cd pipecat
uv sync --extra cartesia --extra daily --extra websocket --extra runner --extra webrtc
A bare uv sync installs the framework but not the extras this example imports: the file pulls in the Daily and FastAPI WebSocket transports and Pipecat’s runner at the top of the file, even when you run it with -t webrtc, so those extras must be present or the import fails with No module named 'daily'. cartesia and webrtc are the parts you actually use.
If you are building your own project instead of working in the Pipecat repo, install the same set:
pip install "pipecat-ai[cartesia,daily,websocket,runner,webrtc]"
Then copy examples/voice/voice-cartesia-turns.py from the repo into your project. The example needs CARTESIA_API_KEY and OPENAI_API_KEY; put both in a .env file next to the script:
CARTESIA_API_KEY=your-key-here
OPENAI_API_KEY=your-key-here
2. Run the agent
uv run examples/voice/voice-cartesia-turns.py -t webrtc
The runner starts a local server and prints a URL, http://localhost:7860 by default. Open it, allow microphone access, and talk. What happens under the hood:
- The browser sends your audio over WebRTC.
CartesiaTurnsSTTServicetranscribes it with Ink 2. The server watches for turn boundaries and emitsturn.start,turn.update,turn.eager_end, andturn.endevents.- The LLM writes a reply.
CartesiaTTSServicestreams the reply through Sonic, and the transport plays it as it arrives.
The interesting part is step 2. Older setups bolt a voice-activity detector onto the audio and guess when a sentence is over; Ink 2’s turn detection runs on the server, which is what makes the next step possible.
3. Pick your voice and your LLM
The example ships with a voice ID and a plain system prompt. Both are yours to change:
tts = CartesiaTTSService(
api_key=os.environ["CARTESIA_API_KEY"],
settings=CartesiaTTSService.Settings(
voice="86e30c1d-714b-4074-a1f2-1cb6b552fb49",
),
)
Pick a voice in the playground and paste its ID here, or generate a new one from a short sample with voice cloning. Because the model speaks the text, keep the system prompt spoken-friendly: the upstream example instructs the LLM to avoid emojis, bullet points, and other formatting nobody can hear.
If you would rather hear a specific Sonic model, pass it in the same settings object. The service currently defaults to sonic-3.6.
4. Cut half a second with eager turn ends
When you finish a sentence, Ink predicts you are done before you have technically finished, and emits turn.eager_end. Pipecat surfaces that as the on_turn_eager_end(service, transcript) event, so you can start generating the reply immediately, then resume if the user keeps talking:
stt = CartesiaTurnsSTTService(
api_key=os.environ["CARTESIA_API_KEY"],
enable_eager_end_of_turn=True,
)
@stt.event_handler("on_turn_eager_end")
async def on_turn_eager_end(service, transcript):
# Ink predicted the user's turn is ending. Start the LLM now.
...
The enable_eager_end_of_turn=True flag matters: it tells Pipecat to act on the server’s early prediction instead of waiting for the committed end of turn. Pipecat publishes a speculative user aggregator example that does the whole dance properly: it starts the LLM on the eager prediction and keeps listening in case you continue. Cartesia’s docs put the saving at around half a second, which is the difference between an agent that answers and one that waits its turn.
If the agent keeps jumping in while you are thinking, tune the thresholds rather than switching eager ends off. The four knobs (and Cartesia’s defaults) are turn_start_threshold (0.8), turn_eager_end_threshold (0.4), turn_end_threshold (0.2), and turn_end_timeout_ms (5600). A patient agent waits longer before betting that you are done, which means raising the quality bar it applies before it interrupts: lower the eager-end threshold toward 0.3 and raise the timeout toward 8000.
A list of names your users will actually say (product names, company jargon) goes in Settings(keyterm=["Pipecat", "Ink 2"]), which nudges the transcription to get them right.
Next steps
- Put the agent on a phone number: rerun the same example with
-t twilioand a public proxy, and the pipeline works over calls unchanged. - If you would rather not own the pipeline at all, build the same agent on Managed Agents, which handles telephony, tools, and knowledge for you.
- Need audio files instead of a live conversation? Batch synthesis is
CartesiaHttpTTSServicefrom the same package.
Related documentation
- Cartesia’s Pipecat integration docs
- Pipecat’s Cartesia STT service and TTS service
- The voice-cartesia-turns example this guide runs
- Speculative user aggregator, the full eager-end pattern
- What makes a voice agent feel real time, for the latency budget behind these choices