Conversational AI design for voice agents

Rene, Kabir Goel 
Conversational AI design for voice agents

Conversational AI design is deciding what an AI agent says, when it says it, and what happens when the conversation goes wrong. Most guides to it are written for chatbots. This one is for voice agents, where the caller can’t scroll back, every pause is audible, and the speech recognizer will occasionally hear “fifteen” when someone said “fifty.”

We build the speech models and the agent platform that voice agents run on, so this guide points to the Cartesia setting that handles each problem. The ideas apply to any stack. The short version: design for the task, write for the ear, and spend most of your effort on the moments when things go sideways. Nobody has ever enjoyed hearing “Sorry, I didn’t catch that” three times in a row.

Voice and chat are different design problems

A lot of conversation design advice carries over from chat. Much of it doesn’t.

ChatVoice
ReadingThe user can reread, skim, and scroll backThe caller hears each word once, in order
LengthA paragraph with a list is fineTwo sentences is a long turn
ChoicesButtons and linksThe caller has to remember the options
SilenceNobody notices a two-second waitA one-second pause sounds like the line dropped
Input errorsTypos are visible to the userRecognition errors are invisible until the agent repeats them
InterruptionsRareConstant, and the agent has to stop talking

Usability research has said this for a while. When Nielsen Norman Group tested voice assistants in 2018, they found the assistants worked well only for simple questions with short answers. LLMs have made agents much better at understanding the question. They haven’t changed how much a person can hold in their head while listening.

Start from the task, not the persona

Before writing a greeting or picking a voice, write down the two or three tasks the agent has to complete and what “done” means for each. “Reschedule an appointment” is done when the new time is in the calendar and the caller has heard it read back. “Answer billing questions” is not a task. “Explain a charge on the last invoice and offer a refund if it’s a duplicate” is.

For each task, list the information the agent needs, where it comes from (the caller, your CRM, a tool), and what the agent may not do. That list becomes the agent’s instructions and its tool descriptions. A persona can come later. A friendly agent that can’t find the order is still the wrong agent.

Write for the ear

Spoken responses need different habits from written ones:

  • One question per turn. “What’s your date of birth and the zip code on the account?” gets half an answer. Ask for one, then the other.
  • Put the answer first. “Your order shipped Tuesday and arrives Friday” beats a sentence that starts with tracking details and gets to the date at the end.
  • Offer few options. A rule of thumb is three at most; by the fourth, callers have forgotten the first. If there are more, ask an open question and let the language model map the answer.
  • No formatting. Markdown, bullet points, and emoji get read aloud or dropped. Tell the model it’s writing for speech.

Numbers, codes, and names need the most care. Sonic reads conventional written forms correctly on its own: prices like $19.99, dates like 04/20/2025, phone numbers like (415) 555-1212. For a confirmation code that must be read character by character, wrap it in a <spell> tag. For a product or place name the model mispronounces, add it to a pronunciation dictionary once instead of fighting it in every prompt. Our prompting tips include a starter system prompt for LLM-written speech that covers these rules; it’s a good first draft of your own.

Test the strings your callers actually hear. An agent that sounds great reading a greeting can still mangle “your balance is $1,204.50, due 11/03.”

Design the timing

In conversation, people hand the floor to each other in about 200 milliseconds (Stivers et al., 2009). An agent that waits a full second to be sure the caller is done sounds slow. One that jumps in at every pause interrupts people who are still thinking.

That decision belongs to turn detection, not to a fixed silence timer. Ink, our speech-to-text model, detects turns from the meaning of what was said as well as the audio. It can emit an early signal when the caller is probably done, so your agent can start preparing a reply, and take it back if the caller keeps going. Our guide to voice activity detection and turn detection goes deeper.

Two more timing decisions are design decisions, not engineering ones:

  • Say something before slow work. If a tool call takes three seconds, silence for three seconds sounds broken. A short line, “Let me pull up that order,” tells the caller the agent heard them. On Managed Agents, tools have a pre_tool_speech setting for this.
  • Decide what can be interrupted. Callers interrupt constantly, and usually the agent should stop and listen. A legal disclosure or a read-back of a payment amount is the exception: finish it, then take the question.

Confirm what matters, and only that

Confirming everything is tedious. Confirming nothing is how an agent books a dentist appointment for the wrong Tuesday. Sort each piece of information by what a mistake would cost:

  • Low cost: the caller’s name in a greeting, the topic of the call. Confirm implicitly by using it: “Sure, I can help with your order.”
  • High cost: dates, amounts, addresses, account numbers, anything that triggers an irreversible action. Read it back and wait for a yes: “That’s Thursday, October 8 at 2:30 PM. Should I book it?”

Recognition errors cluster on the same words: product names, street names, account codes. If you know them in advance, give them to the recognizer as keyterms so it hears them correctly in the first place. Fixing a wrong transcript after the fact is harder than preventing one.

Give every failure a next step

Most of the design work is in the unhappy paths. Plan these before launch:

SituationWhat the agent should do
Didn’t understandRephrase the question, more narrowly. Don’t repeat it word for word.
Didn’t understand twiceOffer a person or another channel. A third “sorry?” is how callers start mashing zero.
SilencePrompt once, then offer to call back or end politely.
Out of scopeSay what the agent can help with and offer a handoff.
Tool failedTell the caller the system is unavailable and give the next step. Never invent an answer.
Caller asks the agent to ignore its instructionsDecline and continue the task. Test this before launch.

Write the two-strikes rule into the agent’s instructions explicitly. LLMs are endlessly patient, and endless patience on a phone call feels like a trap.

Make the handoff easy to reach

Every agent needs a way to reach a person, and the caller should be able to ask for it at any point. On Managed Agents, the transfer_to_number system tool routes the call to a phone number or SIP address, with a plain-language condition for each destination, such as “The caller has a billing or payment question.”

Know which kind of transfer you have. In a cold transfer, the call goes straight to the destination and the agent drops off. In a warm transfer, someone briefs the person picking up before the caller joins. The transfers Cartesia documents today are cold; the SIP trunking docs describe them as SIP REFER transfers. Design for that: have the agent tell the caller who they’re being connected to and why, and if the person picking up needs context, write a short summary to your CRM with a webhook tool before the transfer runs, so nobody has to ask “and what’s this about?”

Test with real calls

A script read aloud by the team that wrote it will pass. Real callers won’t read the script. Before routing traffic:

  • Collect 20 to 50 real recordings or transcripts from the queue you’re automating and run them through the agent.
  • Include the hard ones: background noise, accents, people who change their mind mid-sentence, people who ask for a human in the first five seconds.
  • Score outcomes, not vibes. Managed Agents can grade completed calls with custom metrics, which are LLM judges you define (“Was the caller’s issue resolved?”). It also records time to the first byte of audio on the agent’s first turn of every call.
  • Listen to a sample yourself every week after launch. Transcripts hide tone, long pauses, and talking over the caller.

The AI call center pilot guide turns this into a scorecard for one queue. To build an agent you can test this way, start with Managed Agents and a free account in the Cartesia playground.

FAQs

What is conversational AI design?

Conversational AI design is deciding what an AI agent says, when it says it, what it asks for, and what happens when the conversation goes wrong. It covers the agent's instructions, its wording, how it confirms information, how it recovers from misunderstandings, and when it hands off to a person. For voice agents it also covers timing: when to start speaking, how to handle interruptions, and how to read numbers and codes aloud.

How is designing a voice agent different from designing a chatbot?

A caller can't scroll back, skim, or click a button. Everything has to work in a single pass, in order, and in real time. Voice agents need shorter turns, one question at a time, explicit read-backs for codes and numbers, and a plan for speech recognition errors and interruptions that a text chatbot never sees. Silence matters too: a one-second pause on a call reads as confusion, while the same pause in a chat window goes unnoticed.

What does a conversation designer do?

A conversation designer maps the tasks an agent must complete, writes the prompts and responses, defines how the agent confirms details and recovers from errors, sets the rules for escalating to a human, and tests the result against real conversations. On LLM-based agents, much of the work moves from writing every line to writing instructions, tool descriptions, and test cases that constrain what the model says.

What are the principles of conversational AI design?

Design around the task the caller came to do. Keep each turn short and ask one thing at a time. Confirm only the details that are expensive to get wrong. Give every failure a next step, and never trap someone in a loop. Make the handoff to a person easy to reach. Test with real recordings and real callers, not just a script read by the team that wrote it.

How long should a voice agent take to respond?

In conversation, people take turns with gaps of about 200 milliseconds, so an agent should start answering well under a second after the caller stops. When a tool call will take longer, have the agent say a short line first ("Let me check that order") so the caller knows it heard them.