Voice AIs speak better than they listen, according to Hume's benchmark
Hume's Real World VoiceEQ benchmark evaluates forty voice models. The limits of speech understanding in the face of background noise and accents.
Hume proposes a new benchmark for AI voices, Real World VoiceEQ, which scores over 40 proprietary and open-source models across four areas: automatic speech recognition (ASR), text-to-speech (TTS), speech-to-speech, and spoken language understanding. The suite is built on over 700,000 human evaluations covering various accents, ages, speaking styles, emotions, and acoustic environments, and is divided into separate leaderboards by category.
The company's first finding: there is no single "best" voice model; the field is specializing. Some excel in expressive speech, others in accuracy, while still others capture emotion but struggle to respond naturally. Most importantly, models speak better than they listen: many still reason based on transcription rather than vocal cues like tone, rhythm, hesitation, or emphasis. A confident "yeah" and a hesitant "...yeah" do not call for the same response when dealing with a fraud alert or health advice.
Another point raised: traditional benchmarks overestimate real-world performance, with models shining in clean conditions before stumbling on accented voices, noise, overlapping speakers, and long conversations. Speech-to-speech remains one of the least mature frontiers, with understanding and responding remaining two separate capabilities. From this, Hume makes a case for human evaluation, which it deems irreplaceable, and for a measurement layer blending public benchmarks, private evaluations, and production monitoring. The entire system runs on Kairos, its evaluation platform dedicated to voice.