Voice quality is now a latency problem
Technology
Synthetic speech stopped sounding synthetic a while ago. What still decides whether a voice product works is the round trip: how long a caller waits between finishing a sentence and hearing one back. That is the axis a speech vendor competes on now.
We build voice agents on Amazon Connect and Lex with Bedrock. On that stack, transcription and speech synthesis take one to two seconds before the model is handed a transcript, and responses land in two to three. Every speech vendor gets judged against that clock.
The clock
What a speech vendor is actually competing on
- The model gets whatever the speech stack leaves
On a live call, recognition and synthesis spend their share of the budget before the model sees a transcript. On our own voice agent that is one to two seconds out of a two to three second response. A faster speech model buys back time nothing downstream can.
- Native telephony speech is already paid for
Amazon Connect and Lex handle speech in and out on the stack that already owns the call. Routing audio to a third party adds a network hop and a second vendor to a path measured in hundreds of milliseconds. Worth it when the voice itself is the product.
- A cloned voice needs a consent trail
The technical part of voice cloning is solved. The governance is not: who authorized the likeness, what it is permitted to say, and how it gets revoked when the person behind it leaves or objects. That paperwork is a real cost of the feature and belongs in the plan.
A synthesis vendor earns its place where the voice is the deliverable: narration, dubbing, a branded agent that has to sound like one specific person across thousands of calls. Where the platform already owns speech in and out, the case is harder to make and worth making explicitly.
ElevenLabs has not appeared in a client build we have shipped. Our production voice work runs on Amazon Connect and Lex with Bedrock, where the speech layer comes with the platform and the latency budget is set before anyone picks a voice. Start a conversation.