Neural text-to-speech (neural TTS) is speech synthesis in which a neural network, trained on hours of people talking, generates the audio. Before it, computers either glued together snippets of recorded speech or computed a voice from hand-written rules. Neural TTS learns pronunciation, rhythm, and intonation from data instead, which is why it sounds like a person and not a GPS unit announcing “recalculating.”
If you got here because a settings menu or a cloud console offered you a “neural” voice: pick it. It will sound noticeably more natural. On a cloud bill it usually costs about four times as much as a standard voice, which the pricing section below covers.
If you are choosing a voice model for a product, the more useful question in 2026 is which neural TTS you want. Nearly every voice you’ll compare is neural. They differ in how they generate audio, how fast they start talking, and how they handle the text your product actually sends.
What came before neural TTS
Two older approaches still power the “standard” voices on most cloud platforms.
Concatenative synthesis cuts recordings of a real speaker into small units and splices them back together. Amazon Polly’s documentation describes its standard engine this way: it “concatenates phonemes of recorded speech,” and the seams between units limit how natural it can sound.
Parametric synthesis skips the recordings at playback time. A model predicts acoustic parameters and a signal-processing vocoder turns them into sound. Google Cloud’s voice types page says many of its standard voices use a variation of this. It’s flexible and compact, and it is the source of the slightly buzzy, robotic quality people associate with old text-to-speech.
How neural TTS works
A typical neural TTS system has three stages.
- A text front end normalizes the input. “Dr. Lee, 221B Baker St., $4.50” has to become words a speaker would say, and the model needs to know that “read” in “I read it yesterday” rhymes with “red.”
- An acoustic model predicts what the speech should sound like, usually as a mel spectrogram, a picture of how energy is spread across frequencies over time. This is where the network decides timing, stress, and pitch.
- A neural vocoder turns that spectrogram into a waveform you can play.
The milestones are worth knowing because vendors still name products after them. DeepMind’s WaveNet (2016) generated raw audio one sample at a time and was rated more natural than the best parametric and concatenative systems of its day. Google’s Tacotron 2 (2017) paired a spectrogram predictor with a WaveNet vocoder and scored a mean opinion score of 4.53, against 4.58 for professional recordings. HiFi-GAN (2020) made high-quality vocoding fast enough to run cheaply. Amazon’s neural engine follows this same two-part design, a sequence-to-sequence network that outputs spectrograms followed by a neural vocoder.
The newer generation: speech as tokens
Since about 2023, many systems have dropped the spectrogram. A neural audio codec, such as Meta’s EnCodec, compresses audio into a stream of discrete tokens. A model then predicts those tokens the way a language model predicts words, and the codec decodes them back into sound. Microsoft’s VALL-E showed the payoff: trained on 60,000 hours of English, it could imitate a new speaker from a three-second recording.
This is what cloud providers now sell as “generative” or “HD” voices. Amazon describes its generative engine as a billion-parameter transformer that turns text into speech codes, followed by a decoder that streams them out as audio. These models are better at reading a line the way it’s meant, laughing on cue, and sounding like a conversation. There is a trade-off. Microsoft says its HD voices vary slightly with each output, and Amazon warns that a model update can change how a voice sounds. If you are producing a season of a podcast, that matters. If you are answering phone calls, it mostly doesn’t.
Where state space models fit
Generating speech in real time is a long-sequence problem: a few seconds of audio is thousands of steps. A transformer’s attention looks back over everything generated so far, so each new step gets more expensive as the audio grows. A state space model carries a fixed-size summary of the past forward instead, so each step costs the same.
Cartesia’s founders worked on that idea before Cartesia existed. Albert Gu and Karan Goel co-authored S4, and Gu co-authored Mamba with Tri Dao. Sonic, Cartesia’s TTS model, is built on that architecture family. It streams audio while it generates with sub-90ms model latency, speaks 44 languages, and Sonic 3.6 ranked first on both of Artificial Analysis’s text-to-speech leaderboards when it launched in August 2026. We built it for voice agents, where the caller hears every millisecond of delay.
What “neural” means on a pricing page
Cloud providers use the word as a price tier. Here is what Amazon and Google listed on October 3, 2026, per million characters:
| Provider | Older voices | Neural | Newest tier |
|---|---|---|---|
| Amazon Polly | Standard: $4 | Neural: $16 | Generative: $30; Long-Form: $100 |
| Google Cloud | Standard: $4 | Neural2: $16 | Chirp 3: HD: $30 |
Two things stand out. First, the names don’t line up across vendors, so compare by architecture and listening, not by label. Google now prices its WaveNet voices, the 2016 neural breakthrough, at the same $4 as standard voices. Second, the newest tier costs roughly twice the neural tier. Whether it is worth it depends on your text and your listeners, which you can only find out by testing.
How to tell whether a neural voice is good
Every vendor will tell you its voices sound human. The useful tests are the ones that match your product:
- Send the strings your users will hear: order numbers, email addresses, dates, prices, and names. Neural models fail on these differently from ordinary sentences.
- Listen in context. A line that sounds great alone can be the wrong tone after an apology.
- Measure the path your users experience. For a live conversation, that’s time to first audio over a streaming connection, not how long a finished file takes to render.
- Try long passages and every language you plan to ship. Pacing drifts and accents slip in ways short demos hide.
Our research team wrote a longer piece on this, Is this TTS model good?, which explains why no single score settles the question.
Open-source neural TTS
You can run neural TTS on your own hardware for free. Piper is a fast local engine popular in home automation. Kokoro is an 82-million-parameter model under the Apache license. Both are good choices for reading text aloud on a device without a network connection. They are not built for low-latency streaming at call-center volume, which is the job Sonic does.
To hear the difference yourself, paste a sentence your product actually says into the Cartesia playground. The free plan includes enough credits to try it on your own text.