You know the moment on the news where the anchor throws to a correspondent in another country, and there's that gap. Two seconds of the correspondent smiling gamely into the middle distance while the question crawls up to a satellite and back. Nobody has said anything wrong. It's still deeply uncomfortable, and everyone at home shifts in their seat.
That gap is the single biggest reason AI phone agents feel like AI phone agents. Not the voice. The gap.
So when I sat through Fish Audio's restaurant demo — an agent taking a pizza order — the thing I was watching for wasn't whether it sounded nice. Plenty of things sound nice now. I was watching the clock. It came back in roughly three tenths of a second, every time, and the conversation had an actual rhythm to it.
Then it did something harder. There were people in the background arguing about what to order — talking over each other, changing their minds — and the agent still took the right order off the person who was actually speaking to it. It also switched to the caller's preferred language without any fuss. That is a meaningfully different demo from the usual one-person-in-a-quiet-room performance.
What Fish Audio Actually Shipped
S2.1 Pro is Fish Audio's current state-of-the-art voice model, written up by founder and chief scientist Shijia Liao. The company's own claims: a 61% win rate against the previous generation S2 Pro in head-to-head listening evaluations, roughly 70ms time-to-first-audio on a single request — down from about 100ms in the prior generation — and more than double the throughput under high-concurrency load. On standard API calls they quote around 90ms.
It covers 83 languages through a single model, with no separate endpoints per language, which is presumably what I was hearing when it changed languages mid-demo. It clones a voice from a short reference sample. And Fish Audio has made the same model developers pay for available free via API — model string s2.1-pro-free — with no hard character cap, currently through August 31, 2026, after extending the original deadline twice.
Why Speed Beats Beauty on a Phone Call
Human conversation runs on a startlingly tight clock. We take turns with gaps measured in a couple of hundred milliseconds, and we read anything longer as meaning something — hesitation, confusion, bad news, a bad line. Nobody consciously notices a fast reply. Everybody notices a slow one, and then starts filling the silence themselves, which is the exact moment a caller begins talking over the agent and the whole thing derails.
Which is why a pizza order is a genuinely good test rather than a cute one. Ordering food is fast, interruptive, and full of corrections. Actually, make that a large. No mushrooms. Sorry, what was the total? An agent that needs a beat to think before every reply survives a scripted demo and dies on a Friday night.
So the next time you're comparing vendors on how warm the voice sounds: voice quality is table stakes in 2026. Nobody has ever lost a customer because the agent's vowels weren't lovely enough. They lose them to the pause.
Why My 0.3 Seconds and Their 90ms Are Both True
Here's the part worth understanding, because it explains a discrepancy you'll hit with every vendor you evaluate. Fish Audio quotes about 90 milliseconds. I clocked roughly 300. Neither of us is wrong — we're timing different things.
A voice agent on a live call does three jobs in sequence: it converts the caller's speech to text, runs that through a language model to work out what they want and what to say, then converts the reply back into speech. Time-to-first-audio measures the third leg only — how fast the speaking starts once there's something to say. Everything ahead of it, plus the telephony network, lands in the gap the human actually experiences.
And notice that the most impressive thing in that demo wasn't the speaking leg at all. Pulling one person's order out of a room where several people are talking over each other is the listening leg doing serious work. A beautiful voice model sitting on top of mediocre transcription in a noisy room will confidently read back the wrong order, quickly.
Two caveats before anyone gets carried away. Fish Audio is explicit that the free tier carries no SLA and no latency guarantee — it's best-effort, built for prototyping, not a contractual commitment. And free-tier requests may be retained to improve model quality, which is a decision you want to make deliberately when the audio is your customers' phone calls rather than test sentences.
What to Actually Do With This
Stop grading the voice and start grading the gap. On any demo, insist on interrupting the agent mid-sentence, changing your mind halfway through, and asking something it wasn't expecting. Then do what Fish Audio was brave enough to build into their own demo: make noise. Have someone talk in the background. Order in a second language. That is your Friday evening, and it's precisely what a polished demo reel is designed to avoid.
And when a vendor quotes you a latency figure, ask which leg it refers to. End-to-end — from the moment your caller stops speaking to the moment they hear a reply — is the only number that matches what a human being on the other end of the line actually feels. Everything else is a component spec. Impressive, real, and not the same thing as a conversation that doesn't make somebody shift in their seat.
VENTEEVE