Sonic 3.6 tops a ranking based on human votes
Datapoint has released a dataset containing 315,000 human preferences across 15 text-to-speech models and 300 customer-support scripts, along with the audio samples, individual votes, and methodology behind its Elo ranking.
A voice can pronounce every word correctly and still fail to reassure a customer, read out an order number, or acknowledge a mistake. To compare models in these situations, Datapoint asked evaluators to listen to two versions of the same text and choose the one they preferred.
The company has now released the results as a dataset hosted on Hugging Face. It contains 315,000 human decisions, 4,500 audio files, and 300 English-language scripts designed around customer-support interactions.
Datapoint describes the release as the largest open human-preference dataset dedicated to text-to-speech. That claim is difficult to verify independently without a comprehensive inventory of comparable public datasets. Its scale is nevertheless unusual for an evaluation that provides the audio samples, head-to-head matchups, and individual responses used to produce the ranking.
The protocol compares 15 models in a complete round-robin tournament. Each system is matched against the other 14, producing 105 possible model pairings. Those 105 pairs are evaluated on the same 300 scripts, resulting in 31,500 comparison cells.
Ten valid responses are retained for each cell. Multiplying 31,500 matchups by ten produces the 315,000 published votes. The figure therefore does not represent 315,000 different people or 315,000 audio files.
Each model generated one version of every script, for a total of 4,500 samples. They are stored once as mono FLAC files at 24-bit and 48 kHz, then linked to the relevant comparisons through an identifier. This structure avoids duplicating the same file every time it is compared with a competitor.
Evaluators listened to two samples without being told which models produced them and selected the performance they preferred. Datapoint says it balanced candidate placement so that a model did not consistently appear first or second.
The task was not limited to literal accuracy. Each script included guidance on the expected personality, pacing, and performance criteria. Evaluators could therefore consider clarity, prosody, emotion, pause management, and suitability for the situation.
The 300 scripts are divided into eight categories. Forty-five focus on empathy and de-escalation, another 45 on names and spelling, and 45 on transactional readouts. Forty cover instructions and troubleshooting.
Repairs and disfluencies account for 35 scripts. Brand greetings, policy disclosures, and scheduling or escalation situations contain 30 scripts each.
The selection primarily targets voice agents used by businesses. It includes refunds, appointment scheduling, mandatory disclosures, addresses, case references, and conversations with frustrated customers.
The result should not be treated as a universal text-to-speech ranking. A model preferred when calmly announcing a data breach may perform differently in an audiobook, video game, advertisement, dubbed production, or improvised conversation.
Language further narrows the benchmark’s scope. All scripts are in English. It does not assess French pronunciation, mid-sentence language switching, or the preservation of a consistent voice across multiple languages.
Datapoint initially collected 357,651 responses. The company removed 41,873 responses that its trust mechanism considered invalid, followed by another 778 responses that exceeded the planned quota for certain matchups. That left exactly 315,000 decisions for the September 1, 2026 ranking.
Each published response includes the selected candidate, the time taken to answer, the country reported during the evaluation, and a trust score recorded at the time of participation. The original participant identifiers are not provided. They are replaced with pseudonymous hashes created specifically for this export.
This mechanism allows researchers to examine whether the same person answered consistently without directly revealing the account behind the responses. Datapoint nevertheless prohibits attempts to identify participants by combining the hashes with timestamps and countries.
The trust score accompanies each vote but is not used as a weighting factor in the official ranking. Invalid responses are removed, after which every remaining answer counts equally. Weighted totals are provided separately to show whether the ordering changes when evaluators receive different levels of influence.
The documentation does not fully explain how the trust score is calculated, which thresholds led to exclusion, or how many distinct people participated. It also does not provide a complete demographic breakdown. Datapoint’s claim that it can reach more than 9 million annotators describes the potential reach of its platform, not the population used for this study.
The ranking does not apply Elo directly as a chronological sequence of matches in which scores change after every result. Datapoint first fits a Bradley–Terry statistical model to all valid choices, then converts the estimated strengths to an Elo scale.
Each category receives its own calculation. The overall ranking then combines the eight results with equal weight. A category containing 30 scripts therefore contributes as much to the final score as one containing 45.
This prevents the largest categories from automatically dominating the result, but it also reflects an editorial choice. A service primarily handling transactional readouts might prefer to give that category more weight instead of relying on the overall ranking.
Datapoint uses resampling across groups of scripts to estimate uncertainty. The site consequently displays a possible range of ranks rather than a single fixed position when several models remain difficult to separate.
In the ranking updated on September 3, Cartesia’s Sonic 3.6 leads with an Elo score of 1,103. Speechify’s Simba 3.2 follows at 1,080.
Grok TTS scores 1,037, while Gemini 3.1 Flash TTS reaches 1,033. Both rank ranges cover third and fourth place. The four-point difference is therefore too small to present their ordering as firmly established.
Microsoft’s