The Channel Tax: Why Text Benchmarks Overstate Voice Agents
A benchmark ran 278 identical customer-service tasks in text and again as live phone calls. Voice retained about 79% of the text result. Three of the four named failure modes are conversation mechanics, not intelligence, which changes how you should evaluate any voice vendor.
ScaleVoice
September 1, 2026 · 8 min read
Direct answer
Running identical customer-service tasks in text and in voice shows that voice systems retain roughly 79% of their text performance, up from about 45% when the comparison began. The gap is not caused by weaker reasoning. The named failure modes are speech recognition errors, policy and tool-use mistakes, turn-taking mismanagement, and state tracking across interrupted dialogue, and three of those four are conversation mechanics. A capability demonstrated in text has therefore not been demonstrated for a phone deployment, and evaluations should be run in the voice channel under realistic noise and interruption.
A benchmark published this year ran the experiment this category badly needed and nobody had run cleanly. It took a set of customer-service tasks that already existed as a text evaluation and ran them again as live phone calls. Same 278 tasks. Same tools. Same policies. Same grader. Byte for byte identical, so the only variable left is the channel.
The result is the most useful number in applied voice AI right now: voice systems currently retain about 79% of what the very same systems achieve in text. When the measurement started, that figure was about 45%.
There are two honest ways to read it, and you need both.
The optimistic reading
The gap is closing fast. On the same benchmark the frontier moved from a 30% pass rate in August 2025 to 67% in April 2026, with most of the jump arriving in roughly two months once reasoning-capable voice models landed. Anyone who called voice agents a toy in 2024 has been comprehensively overtaken by the data.
The operational reading
There is still a channel tax of roughly a fifth, and it is not a tax on intelligence.
Look at what the benchmark says actually goes wrong. It names four failure modes:
- Speech recognition errors.
- Policy and tool-use mistakes.
- Turn-taking mismanagement.
- State tracking across interrupted dialogue.
Read that list twice. Three of those four have nothing to do with whether the model understands a service menu or a return policy. They are conversation mechanics: hearing correctly, holding the floor, and remembering where you were when somebody cut you off.
This is not a latency argument and it is not an argument about which vendor is smartest. It is narrower and more awkward than either: a capability that was demonstrated in text has not been demonstrated.
Three things that follow
Ask which channel the claim was measured in
Nearly every agent capability published this year, including multi-step task completion, tool use, policy adherence and handling ambiguity, was established in text, because text is where the benchmarks and the demos live. When a vendor says their agent handles complex multi-step scheduling, that sentence means two different things depending on the channel it was proven in. The honest vendors will tell you which. The unhelpful answer is not usually a lie; it is a claim that quietly assumes full transfer, and the measurement says transfer is about 79%.
The demo never interrupts you
The benchmark deliberately introduces realistic audio conditions, including overlapping speech, background noise and accent variation, and every provider degrades when it does.
Your building is the realistic condition. Real callers talk over the greeting. They say "no, wait, not that car, the other one" halfway through a sentence. There is a compressor running, or a child, or a speakerphone in a parking lot with the engine on. A scripted demo contains none of this, which is why a flawless demo is evidence about the demo and nothing else.
Test state recovery, because almost nobody does
Here is a single evaluation task that costs ten minutes to write and that a surprising number of production systems fail.
The caller starts booking an appointment, changes the vehicle mid-sentence, then asks something unrelated such as how late you are open or whether there is a loaner, and then says "anyway, back to the booking."
A competent human service advisor holds all of that without effort. An agent has to keep a half-built booking alive across two topic changes it did not initiate. That is the fourth failure mode, and it is the first test worth putting in front of any vendor.
What this looks like inside a running deployment
The benchmark described the last two years more accurately than an academic-style evaluation had any right to.
Across the operations behind this article, spanning more than 250 rooftops in four countries, 1.4 million calls and 42,000 booked appointments, almost none of the hard engineering time went into domain knowledge. Teaching a system what a transmission service is, or which advisor handles fleet, or what the loaner policy says, was measured in days.
What took quarters was everything on that four-item list. What the system does when two people speak at once. Whether it can tell a filler word from a real interruption. Whether a booking survives the ninety seconds a caller spends walking out to the driveway to read a vehicle identification number off a windscreen. Whether a mis-heard word gets confirmed back or silently written into a record.
None of that is glamorous and none of it demos well, which is exactly why deployments are won and lost there.
The counter-case, honestly
That 79% is a moving target and it is moving in the right direction, fast. If retention keeps closing at the rate of the last year, the channel tax stops being a strategic consideration and becomes a footnote, and anyone who built an entire buying process around it will look overcautious within eighteen months.
The benchmark's tasks are also retail, airline and telecom flows, not a service drive at half past seven on a Monday. Transfer from one to the other is an assumption, not a measurement, and the figure should not be quoted as though it applies precisely to any specific deployment. It does not.
But you are buying this year, on this year's number, and this year a fifth is a fifth.
A prediction with a date
By the end of 2027, "we tested it in text" will be an unacceptable answer in a voice procurement conversation, the way "it works on my machine" stopped being acceptable in software somewhere around 2015.
The benchmark that made the comparison possible is nine months old. Buying norms lag measurement by a couple of years, and then they flip quickly.
Next step
See how your first workflow could work in a demo.
Share the call source, booking rules, systems you use, and when your team should step in. ScaleVoice will show how the first workflow can turn that demand into measurable booked outcomes.
Book a demoRelated pages
FAQ
Questions to consider before your first workflow
What is the channel tax in voice AI?
It is the measured performance gap between the same agent completing the same tasks in text versus over a live phone call. Current benchmarking puts voice retention at roughly 79% of text performance, meaning about a fifth of demonstrated capability does not survive the move to speech.
Why do voice agents fail if the underlying model is capable?
Because most named failure modes are conversational rather than cognitive. Speech recognition errors, turn-taking mismanagement and state tracking across interrupted dialogue all occur even when the model knows the correct answer. Only policy and tool-use mistakes touch reasoning directly.
How should I evaluate a voice agent vendor?
Run the evaluation in the voice channel, under realistic conditions including background noise, accents and interruptions, using your own scenarios. Ask explicitly which channel any capability claim was measured in, and include at least one task where the caller changes subject mid-transaction and then returns to it.
What is a good interruption test for a voice agent?
Have the caller begin a booking, change the vehicle mid-sentence, ask an unrelated question such as opening hours or loaner availability, then return to the booking. The agent must keep the partially completed booking intact across two topic changes it did not initiate.