Customer stories

ActOnCue's AI scene readers moved from ElevenLabs to Sonic 3

Of ActOnCue's speech-to-speech usage runs on Cartesia

93.3%

Users switching from ElevenLabs to Cartesia, against those switching back

3.9% vs 2%

George Rybintsev built ActOnCue for actors recording auditions at home. An actor uploads a script, picks the character they are playing, and rehearses against AI voices reading the other parts. Cartesia is the default provider for those voices, and Sonic 3 is what moved ActOnCue off ElevenLabs.

The challenge: the solo job that needs a second person

Self-tapes were a pandemic workaround that stuck. Casting still runs on them, which leaves actors to produce their own auditions.

A self-tape is the only solo endeavor that requires a crew. One person can run the camera and read the second part, but that doubles the effort, the time, and the editing. In practice, a self-tape means calling your roommate at 10pm and asking them to play Juliette.

ActOnCue takes the second person out of the equation without handing the work back to the actor. Drop in a script, select your character, and start rehearsing. The platform finds the remaining characters and assigns voices to them automatically, and any of those voices can be changed after the first run. An actor should not have to configure an AI system before an audition.

You just want to go from A to B really fast and really efficiently.

George RybintsevFounder, ActOnCue

Since launching, ActOnCue has grown into the rest of the acting workflow: line learning, rehearsal, and preparation for different acting disciplines. Voice became the most important part of the experience.

Why ActOnCue chose Cartesia

A scene reader has an unusual job. It has to be present enough to react to and plain enough to ignore. The actor needs something to play against, not a performance. The moment the reader starts making choices of its own, the actor starts shaping their lines as responses to those choices, and the self-tape stops being theirs.

You don’t want it over-dramatized during self tapes. You want it to be very consistent. Clean, and make sure that the voice, the reader’s voice, doesn’t take away from the self tape.

George RybintsevFounder, ActOnCue

ActOnCue started on ElevenLabs and found the voices too dramatic for the job. They performed the lines and competed with the actor instead of supporting them. ElevenLabs also cost more, and its roadmap was visibly moving toward consumer products rather than the developer surface ActOnCue needed.

Sonic 3 was the turning point: a conversational cadence and human delivery that does not overpower the performance. Cartesia brought ActOnCue into its startups program as a Cartesia Startup Grant recipient, with 12 months of free access to the Scale plan.

The solution: the default on a multi-vendor platform

ActOnCue still offers several voice providers, and Cartesia is the one it assigns by default. That choice came out of a close reading of what a rehearsal actually needs, and it is not binding: actors can move between providers at any point.

What the reader stopped costing them is the result the team cares about most. A voice that holds a neutral register on its own does not need managing, tuning, or apologizing for, which leaves ActOnCue free to build the rest of the actor’s workflow and leaves the actor with the only performance that matters in the room.

The results: actors keep the default

  • 93.3% of ActOnCue’s speech-to-speech usage runs on Cartesia, on a platform where actors can pick any provider they want
  • 3.9% of users who start on an ElevenLabs voice later switch to Cartesia, against about 2% who go the other way, making a switch nearly twice as likely to land on Cartesia as to leave it

Looking ahead

ActOnCue’s next release is a full overhaul built around giving actors a linear workflow. Auto-assigned voices were the first move in that direction. The overhaul is the rest of it: pulling every remaining piece of machinery off the actor’s path and putting it behind the product. The simpler that path gets, the more the voice underneath it has to carry.

Architecting AI that learns and interacts like humans.

Status