What is neural TTS? How neural voices work

Rene, Kabir Goel 
What is neural TTS? How neural voices work

Neural text-to-speech (neural TTS) is speech synthesis in which a neural network, trained on hours of people talking, generates the audio. Before it, computers either glued together snippets of recorded speech or computed a voice from hand-written rules. Neural TTS learns pronunciation, rhythm, and intonation from data instead, which is why it sounds like a person and not a GPS unit announcing “recalculating.”

If you got here because a settings menu or a cloud console offered you a “neural” voice: pick it. It will sound noticeably more natural. On a cloud bill it usually costs about four times as much as a standard voice, which the pricing section below covers.

If you are choosing a voice model for a product, the more useful question in 2026 is which neural TTS you want. Nearly every voice you’ll compare is neural. They differ in how they generate audio, how fast they start talking, and how they handle the text your product actually sends.

What came before neural TTS

Two older approaches still power the “standard” voices on most cloud platforms.

Concatenative synthesis cuts recordings of a real speaker into small units and splices them back together. Amazon Polly’s documentation describes its standard engine this way: it “concatenates phonemes of recorded speech,” and the seams between units limit how natural it can sound.

Parametric synthesis skips the recordings at playback time. A model predicts acoustic parameters and a signal-processing vocoder turns them into sound. Google Cloud’s voice types page says many of its standard voices use a variation of this. It’s flexible and compact, and it is the source of the slightly buzzy, robotic quality people associate with old text-to-speech.

How neural TTS works

A typical neural TTS system has three stages.

  1. A text front end normalizes the input. “Dr. Lee, 221B Baker St., $4.50” has to become words a speaker would say, and the model needs to know that “read” in “I read it yesterday” rhymes with “red.”
  2. An acoustic model predicts what the speech should sound like, usually as a mel spectrogram, a picture of how energy is spread across frequencies over time. This is where the network decides timing, stress, and pitch.
  3. A neural vocoder turns that spectrogram into a waveform you can play.

The milestones are worth knowing because vendors still name products after them. DeepMind’s WaveNet (2016) generated raw audio one sample at a time and was rated more natural than the best parametric and concatenative systems of its day. Google’s Tacotron 2 (2017) paired a spectrogram predictor with a WaveNet vocoder and scored a mean opinion score of 4.53, against 4.58 for professional recordings. HiFi-GAN (2020) made high-quality vocoding fast enough to run cheaply. Amazon’s neural engine follows this same two-part design, a sequence-to-sequence network that outputs spectrograms followed by a neural vocoder.

The newer generation: speech as tokens

Since about 2023, many systems have dropped the spectrogram. A neural audio codec, such as Meta’s EnCodec, compresses audio into a stream of discrete tokens. A model then predicts those tokens the way a language model predicts words, and the codec decodes them back into sound. Microsoft’s VALL-E showed the payoff: trained on 60,000 hours of English, it could imitate a new speaker from a three-second recording.

This is what cloud providers now sell as “generative” or “HD” voices. Amazon describes its generative engine as a billion-parameter transformer that turns text into speech codes, followed by a decoder that streams them out as audio. These models are better at reading a line the way it’s meant, laughing on cue, and sounding like a conversation. There is a trade-off. Microsoft says its HD voices vary slightly with each output, and Amazon warns that a model update can change how a voice sounds. If you are producing a season of a podcast, that matters. If you are answering phone calls, it mostly doesn’t.

Where state space models fit

Generating speech in real time is a long-sequence problem: a few seconds of audio is thousands of steps. A transformer’s attention looks back over everything generated so far, so each new step gets more expensive as the audio grows. A state space model carries a fixed-size summary of the past forward instead, so each step costs the same.

Cartesia’s founders worked on that idea before Cartesia existed. Albert Gu and Karan Goel co-authored S4, and Gu co-authored Mamba with Tri Dao. Sonic, Cartesia’s TTS model, is built on that architecture family. It streams audio while it generates with sub-90ms model latency, speaks 44 languages, and Sonic 3.6 ranked first on both of Artificial Analysis’s text-to-speech leaderboards when it launched in August 2026. We built it for voice agents, where the caller hears every millisecond of delay.

What “neural” means on a pricing page

Cloud providers use the word as a price tier. Here is what Amazon and Google listed on October 3, 2026, per million characters:

ProviderOlder voicesNeuralNewest tier
Amazon PollyStandard: $4Neural: $16Generative: $30; Long-Form: $100
Google CloudStandard: $4Neural2: $16Chirp 3: HD: $30

Two things stand out. First, the names don’t line up across vendors, so compare by architecture and listening, not by label. Google now prices its WaveNet voices, the 2016 neural breakthrough, at the same $4 as standard voices. Second, the newest tier costs roughly twice the neural tier. Whether it is worth it depends on your text and your listeners, which you can only find out by testing.

How to tell whether a neural voice is good

Every vendor will tell you its voices sound human. The useful tests are the ones that match your product:

  • Send the strings your users will hear: order numbers, email addresses, dates, prices, and names. Neural models fail on these differently from ordinary sentences.
  • Listen in context. A line that sounds great alone can be the wrong tone after an apology.
  • Measure the path your users experience. For a live conversation, that’s time to first audio over a streaming connection, not how long a finished file takes to render.
  • Try long passages and every language you plan to ship. Pacing drifts and accents slip in ways short demos hide.

Our research team wrote a longer piece on this, Is this TTS model good?, which explains why no single score settles the question.

Open-source neural TTS

You can run neural TTS on your own hardware for free. Piper is a fast local engine popular in home automation. Kokoro is an 82-million-parameter model under the Apache license. Both are good choices for reading text aloud on a device without a network connection. They are not built for low-latency streaming at call-center volume, which is the job Sonic does.

To hear the difference yourself, paste a sentence your product actually says into the Cartesia playground. The free plan includes enough credits to try it on your own text.

FAQs

What is neural TTS?

Neural text-to-speech (neural TTS) is speech synthesis in which a neural network, trained on recordings of people talking, generates the audio. Older systems either stitched together clips of recorded speech or ran hand-built signal processing. Neural TTS learns pronunciation, timing, and intonation from data, which is why it sounds much closer to a person.

What is the difference between neural and standard TTS voices?

On cloud services, "standard" usually means an older concatenative or parametric voice and "neural" means a voice generated by a neural network. Amazon Polly's documentation, for example, says its standard voices concatenate phonemes of recorded speech, while its neural engine predicts spectrograms with a neural network and turns them into audio with a neural vocoder. Neural voices sound more natural and cost more: on October 3, 2026, Polly listed $4 per million characters for standard voices and $16 for neural.

Is neural TTS the same as AI voice?

Mostly, yes. Every modern AI voice generator is a neural TTS model. The newest systems, sometimes called generative voices, are still neural networks; they just predict compressed audio tokens the way a language model predicts words, instead of predicting a spectrogram.

Can neural TTS run in real time?

Yes, if the model streams. A streaming model starts sending audio before it has finished generating the sentence, so playback can begin almost immediately. Cartesia's Sonic, for example, has sub-90ms model latency and streams audio while it generates, which is what voice agents on the phone need.

How much does neural TTS cost?

It depends on the provider and the tier. On October 3, 2026, Amazon Polly listed $16 per million characters for neural voices and $30 for generative voices, and Google Cloud listed $16 per million characters for Neural2 and $30 for Chirp 3 HD. Open-source models such as Piper and Kokoro are free to run on your own hardware. Cartesia's prices are on its pricing page, with a free plan to start.