AI Frontiers, part 41: Voice interfaces β ASR, TTS, and the latency budget
Part 41from the AI Frontiers series · 65 parts in all
Two of the three components of a voice interface are essentially solved. Speech recognition, after a decade of architectural convergence, works well enough on clean audio that word error rates are no longer the interesting number. Speech synthesis, after a similar convergence, produces audio that is difficult to distinguish from a recording in short utterances. And yet the experience of talking to software is still markedly worse than talking to a person, for reasons that have almost nothing to do with either of those models.
The gap is turn-taking. Conventions about when a speaker is finished, how long a pause is polite, how to interrupt, how to signal that you are still thinking β those are what make conversation feel effortless, and they are the parts a pipeline gets wrong. This entry is about the components, the latency arithmetic, and why the latency budget rather than the accuracy is what distinguishes a usable voice interface.
Recognition: what converged, and what still breaks
Automatic speech recognition moved through a recognizable arc. Hidden Markov models with hand-engineered features gave way to neural acoustic models, then to end-to-end architectures β first connectionist temporal classification, then the recurrent neural network transducer (Graves), which remains the dominant production architecture because it streams naturally. Self-supervised pretraining (wav2vec 2.0, Baevski et al.) made it possible to learn from raw audio without transcripts and then fine-tune on a small labeled set, which is what allowed good recognition for languages with little annotated data. Whisper (Radford et al.) took the opposite approach β 680,000 hours of weakly supervised multilingual audio β and produced a single model that is robust across accents, noise and languages.
Whisper's influence was disproportionate to its novelty because it was released as weights and code with a permissive license, which made high-quality transcription free and immediate. Within months it was inside meeting tools, note applications, video editors and the assistant ecosystem, and the volume of speech that suddenly had transcripts attached changes what gets built next.
The failures that remain are consistent and worth knowing. Recognition degrades badly on overlapping speech, which is the normal condition of a conversation between two people. It handles rare proper nouns poorly, which matters most for exactly the domains where transcription has value β medicine, law, engineering β because those vocabularies are full of unusual terms. It hallucinates on silence, producing plausible sentences where there was no speech, which is a data-quality disaster for anyone building a transcript corpus and a particular instance of the problem in part 26: the model continues what it expects rather than reporting the absence of an input. And diarization β deciding who spoke when β is a separate model with its own error rate that compounds with the recognition error rather than adding to it quietly.
Synthesis: from concatenation to codecs
Text-to-speech followed a parallel path with a different endpoint. Concatenative synthesis, which splices recorded fragments, was replaced by parametric neural models β Tacotron 2 (Shen et al.) for the acoustic model and WaveNet-style vocoders for the waveform β then by non-autoregressive architectures that generate the whole spectrogram at once (FastSpeech 2, Ren et al.) and models that generate waveform and spectrogram jointly (VITS, Kim et al.). The practical effect was that synthesis became fast enough for real-time use on modest hardware.
The more consequential shift was to neural audio codecs (SoundStream, Zeghidour et al.; EnCodec, DΓ©fossez et al.) and the language models built on top of them. If audio can be tokenized into a discrete sequence, then a speech model is a sequence model, and every technique in this series applies β autoregressive generation, prompting, instruction tuning, and eventually the efficiency work on caching and quantization. AudioLM (Borsos et al.) established the framing; VALL-E (Wang et al.) demonstrated zero-shot voice cloning from a three-second sample, which is a capability with consequences well beyond user interfaces. The choice between a pipeline of specialized components and a single end-to-end model became a genuine architectural decision rather than a settled one.
Two-sided problems came with the capability. Voice cloning made impersonation trivial, and the identity-verification systems that banks and call centers relied on became simultaneously more valuable and less trustworthy. The synthesis of consent β who authorized this voice, for what, for how long β is a governance question the technology produced faster than the industry answered it.
The latency budget, which is where interfaces live or die
Telephony engineering solved this problem decades ago, and the number is still the right one: around 150 milliseconds of one-way delay is where conversation begins to feel unnatural, and beyond roughly 300 milliseconds the turn-taking conventions break down entirely (International Telecommunication Union). A pipeline that transcribes, reasons and then synthesizes spends its budget in four places, and the sum is what the user experiences.
Endpointing. Deciding the speaker has finished. Too eager and you interrupt people mid-thought; too patient and every exchange carries a dead pause that makes the system feel slow. This is the component that most distinguishes a good voice interface and it is the one most often implemented as a fixed silence threshold, because it is easy and it demos acceptably.
Recognition. Processed incrementally while the user speaks, so that only the final segment is outstanding at the moment of endpointing. A non-streaming model that must process the entire utterance after the fact adds its full inference time to the critical path.
The language model. The largest and most variable term, and the one this series has spent the most time on. Time to first token is what matters, not total generation time, because synthesis can begin on the first clause. Streaming output is therefore not a nicety; without it, a two-second generation becomes a two-second silence.
Synthesis. Also incremental, generating audio as text arrives, which is why a non-streaming vocoder destroys the benefit of a streaming language model.
The arithmetic explains a pattern I have seen repeatedly. A system with an excellent recognition model and an excellent synthesis model, wired together without streaming at any stage, feels noticeably worse than a system with mediocre models and streaming throughout. The models dominate the papers and the plumbing dominates the experience.
Speech as a different interface, not a cheaper one
Voice is usually described as an accessibility feature or a convenience, and treating it as a substitute for typing misses what changes. Speech is fast, low-precision and full of redundancy; text is slow, precise and compressed. Interfaces that work by treating speech as transcribed text therefore throw away the properties that make speech useful and inherit the properties that make text unforgiving.
Three differences are worth designing for. Speech carries prosody β hesitation, emphasis, uncertainty β and an interface that discards it loses the signal a listener uses most. Speech is ephemeral, so any answer the user must remember or act on needs a visible artifact; the good voice interfaces in this period all wrote something down. And speech is social, which means the conventions of interruption and turn-taking are not decoration but the interface itself, and violating them produces a system people describe as rude rather than slow.
There is also a security property worth naming. Voice is now trivially synthesizable in anyone's voice from a short sample, so any identity claim based on voice alone is no longer evidence. Systems that authenticate by voice need a second factor, and systems that act on voice commands need to consider who else in the room can speak.
The long tail: dialects, code-switching and noise
A single word error rate is an average over a distribution that is almost never the distribution you actually serve, and the gap is systematic rather than random.
Accent and dialect variation produces error rates that differ by large multiples between groups, and the pattern is consistent with the training data: a model trained predominantly on a particular variety of a language performs worst on the varieties it has seen least. Code-switching β alternating between two languages within a sentence, which is normal behavior for bilingual speakers and extremely common in some regions β is handled poorly by models with a single-language assumption, and it is frequently the case that matters most for customer-facing deployment.
Acoustic conditions matter as much as linguistic ones. Far-field audio from a room microphone has reverberation and competing sources; overlapping speech breaks an assumption every diarization system makes; a noisy vehicle cabin combines both. Telephony-band audio, still the format for a great deal of customer interaction, discards the high frequencies that disambiguate certain consonants.
The measurement implication is that evaluation must be disaggregated. A system reporting six percent overall error and fourteen percent on the accent group that accounts for a third of its users has a product problem that an average hides, and the fix β more data from the failing group, or a different acoustic front end β is only discoverable through the disaggregation. This is the same argument the data-curation entry makes about filters, arriving from the acoustic side.
One number summarizes streaming feasibility: the real-time factor, meaning the ratio of processing time to audio duration. A system needs to stay comfortably below one to keep up, and the margin is what determines how much buffering is required before a first hypothesis can be emitted. This is the acoustic equivalent of time-to-first-token, and it is why a batch transcription system with excellent accuracy is often a poor foundation for a live interface regardless of how good its model is.
What to measure, and why error rate is not enough
Word error rate is the field's standard metric and it is a poor predictor of whether a voice interface works, for reasons worth spelling out.
It treats all words as equally important, so dropping an article and mangling a product code count the same. It is computed on transcripts with conventions for punctuation, casing and disfluencies that differ between systems, which means published comparisons are frequently not comparing the same thing. And it is upstream of everything the user experiences: a system can have an excellent recognition model and still frustrate people because the synthesized voice answers before they finished speaking.
The metrics that predict usability are these. Latency percentiles rather than an average, because the ninety-fifth percentile is what a user notices and the mean hides it. Barge-in success rate β whether the system stops speaking when interrupted β which is the single most telling behavior of a conversational system and is almost never reported. Downstream task accuracy, which measures whether the intent survived the transcription rather than whether the words were reproduced. Semantic error rate, which weights errors by whether they change meaning. Disaggregated accuracy by accent, language and acoustic condition. And for synthesis, listener ratings of naturalness, which remain the only measure that correlates well with how people describe the experience.
The pattern is the same as everywhere else in this series. The headline metric is easy to compute and weakly related to the outcome; the metrics that matter require defining what the outcome is. Teams that define it ship voice interfaces people use. Teams that track error rate ship demonstrations.
Where the architecture is going
The trajectory is the one part 10 described: collapsed pipelines. Instead of transcribe, reason and synthesize as three models passing strings, one model consumes audio tokens and emits audio tokens, with text as an intermediate that may or may not be materialized. The advantage is that prosody, emotion and timing survive the process instead of being stripped in the first stage and re-invented by the last.
The cost is debuggability and controllability. A cascaded pipeline has inspection points where you can read the transcript, correct it, filter it, and log it β which is how every compliance and safety requirement in this domain gets satisfied. An end-to-end speech model has none, and giving it transcript-level controls means adding a transcriber back. The likely outcome is a hybrid: end-to-end for the conversational path where latency and naturalness dominate, and cascaded where the transcript is the product. Choosing between them is a decision about what you are actually selling, which is the same conclusion the rest of this series keeps arriving at.
Works Cited
Baevski, Alexei, et al. "wav2vec 2.0: A Framework for Self-Supervised Learning of Speech Representations." arXiv, 2020, arxiv.org/abs/2006.11477. Accessed 11 July 2024.
Borsos, ZalΓ‘n, et al. "AudioLM: A Language Modeling Approach to Audio Generation." arXiv, 2022, arxiv.org/abs/2209.03143. Accessed 11 July 2024.
DΓ©fossez, Alexandre, et al. "High Fidelity Neural Audio Compression." arXiv, 2022, arxiv.org/abs/2210.13438. Accessed 11 July 2024.
Graves, Alex. "Sequence Transduction with Recurrent Neural Networks." arXiv, 2012, arxiv.org/abs/1211.3711. Accessed 11 July 2024.
Gulati, Anmol, et al. "Conformer: Convolution-Augmented Transformer for Speech Recognition." arXiv, 2020, arxiv.org/abs/2005.08100. Accessed 11 July 2024.
International Telecommunication Union. "ITU-T G.114: One-Way Transmission Time." ITU-T Recommendation, 2003. Accessed 11 July 2024.
Kim, Jaehyeon, et al. "Conditional Variational Autoencoder with Adversarial Learning for End-to-End Text-to-Speech." arXiv, 2021, arxiv.org/abs/2106.06103. Accessed 11 July 2024.
Kong, Jungil, et al. "HiFi-GAN: Generative Adversarial Networks for Efficient and High Fidelity Speech Synthesis." arXiv, 2020, arxiv.org/abs/2010.05646. Accessed 11 July 2024.
Radford, Alec, et al. "Robust Speech Recognition via Large-Scale Weak Supervision." arXiv, 2022, arxiv.org/abs/2212.04356. Accessed 11 July 2024.
Ren, Yi, et al. "FastSpeech 2: Fast and High-Quality End-to-End Text to Speech." arXiv, 2020, arxiv.org/abs/2006.04558. Accessed 11 July 2024.
Shen, Jonathan, et al. "Natural TTS Synthesis by Conditioning WaveNet on Mel Spectrogram Predictions." arXiv, 2017, arxiv.org/abs/1712.05884. Accessed 11 July 2024.
Wang, Chengyi, et al. "Neural Codec Language Models Are Zero-Shot Text to Speech Synthesizers." arXiv, 2023, arxiv.org/abs/2301.02111. Accessed 11 July 2024.
Zeghidour, Neil, et al. "SoundStream: An End-to-End Neural Audio Codec." arXiv, 2021, arxiv.org/abs/2107.03312. Accessed 11 July 2024.