Text to Speech Benchmarking Methodology
Scope & Background
Artificial Analysis performs benchmarking on Text to Speech models, primarily those delivered via serverless API endpoints. This page describes our Text to Speech benchmarking methodology, including both our quality benchmarking and performance benchmarking. Speed and price metrics require a serverless endpoint. Models without one may still be included in the Arenas for quality comparison only.
Across both our performance benchmarking and the Arenas, our focus is reflecting the end-user experience of the serverless APIs. We use the standard implementation of provider APIs as suggested by each provider's documentation.
Key Metrics
We use the following metrics to track quality, performance and price for Text to Speech models.
Quality Elo
Relative Elo scores of the models are determined by responses from users in the Artificial Analysis Text to Speech Arena. Rankings are computed from pairwise votes using a Bradley-Terry model fit by maximum likelihood. Ratings are anchored by fixing one model per arena at 1000, with other scores reported relative to it on the standard Elo scale. See How the Arena Works below for more detail on how comparisons are presented and voted on.
Matchup sampling favours models that need comparisons, such as recent entrants and models that do not yet have a rating, so new models accumulate comparisons quickly. A cap limits the share of a model's comparisons that any single opponent can receive, and at estimation time battles from frequently repeated pairings are down-weighted, so no single matchup dominates a model's rating. In our experiments, rankings on both the Provider Voice and Controlled Voice Arenas were not sensitive to the strength of this weighting.
Each board is fit independently: the overall rating uses all eligible votes, while category, accent and language boards are separate fits on their own subsets rather than aggregations of the overall result. All eligible votes in a board's history count equally, with no time weighting or recency decay. Ratings are recomputed multiple times daily from the full vote history and reviewed before publication. We display 95% confidence intervals and rank ranges, and models are ordered by unrounded Elo. Some models may not be shown due to not yet having enough votes.
Pronunciation Robustness
The Pronunciation Robustness benchmark measures whether Text to Speech models pronounce challenging text correctly. It complements Arena preference results by scoring every model on pronunciation of a fixed set of challenge sentences, judged by human reviewers for accepted readings.
Every model reads the identical set of sentences with a single pinned US English voice. Each sentence carries one or more highlighted spans with accepted pronunciations, including variants valid in US and UK English. Clips are served to reviewers in random order without model names. Prompts are sent to each model as written, with no normalization on our side. Where a model offers a text normalization setting, we use its default.
The set covers four categories of pronunciation challenges:
- Contextually Appropriate (Contextual Disambiguation): words spelled the same but read differently depending on context, e.g. a wound that is bandaged vs. a bandage that is wound.
- Expanding Shorthand (Text Normalization): numbers, dates, units and notation read out naturally, e.g. 6'2", Chapter XVII, or 1 tsp of sugar.
- Preserving Exact Sequences (Sequence Fidelity): codes, paths, emails and identifiers spoken exactly, e.g. .env.local or a.chen@ucsf.edu.
- Standalone Terms (Term Pronunciation): brand, place and technical names pronounced correctly, e.g. Arkansas, façade, or genre.
Three or more independent third-party reviewers judge over 95% of clips for each published model. For each highlighted span they answer whether it was pronounced naturally and correctly with either of Yes, No, or I could not tell. A model's score is the share of span judgements answered Yes, out of Yes plus No. "I could not tell" answers are excluded from the denominator. Category scores use the same formula over that category's spans. Error bars show 95% confidence intervals, computed with span-clustered standard errors because each span is judged by several reviewers.
We publish results for a model once at least 95% of the set's samples are covered by three or more approved reviewers.
See the Pronunciation Robustness Overview for an example of the rating card reviewers see and example clips for every model on the benchmark.
Price per 1M Characters
All TTS models are reported as price per 1M characters of input text. Where providers do not price per character directly, we derive an equivalent:
- Per character: Listed rate used directly
- Subscription plans: Effective rate from the lowest-cost plan, annual where available, that includes at least 1M characters, assuming 80% utilization
- Token-based: Derived using batch pricing, converting characters to input tokens and estimating audio output duration
- Output duration: Converted using an assumed speaking rate of 150 words per minute, approximately 825 characters per minute
- Per byte: 1 UTF-8 byte is approximately 1 character for English text
- Inference time: Estimated from benchmark runs using ~25 texts of ~500 characters
- Open weights: Listed without price, unless the model is also offered through a paid API, in which case that rate is used
- Speech to Speech models: Derived from the cost to generate speech for 20 fixed TTS arena prompts, based on reported text input, output audio, output text, and separately exposed reasoning/thinking tokens. Output text tokens are included when they are reported as part of the audio response. We exclude caching effects because prompts are run independently as single-turn requests, so their impact on pricing is negligible.
Throughput
Median characters of input text the provider converts to speech per second, calculated over the past 3 days of measurements. For each generation we divide the prompt's character count by the time taken to generate and receive the clip.
The measurement includes downloading the audio clip from the provider where a URL is provided rather than an audio response. This is to reflect the end-user experience of receiving a generated audio clip, and as URLs can be generated prior to audio completion. Audio clips are generated at batch size of 1 where relevant.
Benchmarking is conducted 4 times per day. For each benchmarking evaluation we select a single random voice for each model. A unique prompt of ~500 characters is used for each generation.
How the Arena Works
Each comparison presents two audio clips generated from the same text prompt. In the Provider Voice Arena both clips use voices from the same gender and accent category; in the Controlled Voice Arena both use the same cloned reference voice. Model names are hidden until after voting, presentation order is randomized, and listeners must play a minimum portion of each clip before voting is enabled. Votes are a binary choice with no tie option. Prompts span four categories (Customer Service, Entertainment, Knowledge Sharing, and Assistants), including challenge material such as alphanumeric strings and numeric sequences, and the prompt set is refreshed over time.
Audio samples are generated once per model and standardized before entering the arena, so votes for a model are cast on a consistent set of samples that can be checked and normalized before going live. Clips are converted to a consistent playback format and loudness-normalized to a common target, so that differences in volume or encoding between providers do not influence votes.
Controlled Voice Arena
The Controlled Voice Arena uses voice cloning to standardize the voices used to evaluate each Text to Speech model. By comparing models using the same reference voices, it separates listener preference for a particular voice from broader aspects of model quality, such as speech naturalness, audio quality, pronunciation, pacing, and prosody. This enables a more like-for-like comparison of the underlying Text to Speech models.
For the Controlled Voice Arena, the English reference set consists of 8 professionally recorded voices: 2 US female, 2 US male, 2 UK female, and 2 UK male. Where a model accepts the full reference recording, including models that crop it natively, we provide the full recording. Where a model limits the length of reference audio, we instead provide a shorter version, typically 5 to 10 or 10 to 15 seconds, cut at a natural pause.
Beyond English, the Controlled Voice Arena also evaluates models in Spanish, French, German, Portuguese, Japanese, Hindi, Mandarin Chinese, Arabic, and Vietnamese. For each non-English language, the reference set includes one professionally recorded male voice and one professionally recorded female voice, both recorded by native speakers. Voice cloning is also used to standardize voices within each language, while preference votes are collected from raters screened to ensure the evaluated language is their first language. Each model is evaluated only in the languages it supports, and a language-specific leaderboard is published once models have accumulated enough comparisons to produce stable Elo ratings.
The Controlled Voice Arena complements our Provider Voice Arena, where each model is evaluated using a representative set of its own publicly available voices.
Select a controlled voice below to listen to its reference recording.
Provider Voice Arena
The Provider Voice Arena evaluates each Text to Speech model using a representative set of the voices made available by that model's provider. This complements the Controlled Voice Arena, where models are compared using the same cloned reference voices. Provider voices reflect how each model is typically experienced by users through its own public API, product interface, or documentation.
For each model, we select multiple provider-offered voices to make the comparison representative and fair. We select 2 voices for each combination of male and female, and US and UK accents, for 8 voices in total. Where a gender and accent combination is not available, we exclude that combination from evaluation in the Provider Voice Arena.
Voices are selected based on their prominence in provider interfaces and documentation, excluding voices that are not neutral in nature (for example, highly stylized regional accents). Model creators may also request that we use specific voices where many are available. Where provider voices are not provided, as is typically the case for open source models, we use voice clips from professional voice actors as source files for generating speech (see the Controlled Voice Arena above for samples).
Select a model creator below to see the provider voices currently used for each of their models.
Voter Quality
Most of the ratings we publish come from a paid panel with strong track records on comparable projects, covering both arena preference votes and Pronunciation Robustness reviews. For the English arenas and for Pronunciation Robustness, panel members are screened for native English speakers resident in the US or UK. For each non-English language in the Controlled Voice Arena, panel members are screened on first language, with raters residing in their country of origin where possible, so that raters are native speakers of the language they evaluate.
Pronunciation Robustness sessions include attention-check clips with known answers, and reviewers who miss any answered check are excluded entirely. Arena votes pass automated quality screening before entering rankings. Paid raters never see model names. We compare rankings from the paid and public pools as an ongoing consistency check.
Model and Provider Inclusion Criteria
Our objective is to analyze and compare popular and high-performing Text to Speech models and providers to support end-users in choosing which to use. As such, we apply an 'industry significance' and competitive performance test to evaluate the inclusion of new models and providers. We are in the process of refining these criteria and welcome any feedback and suggestions. To suggest models or providers, please contact us via the contact page. When a provider publicly releases a material update to a model, we list it as a new dated variant, regenerate all of its arena samples, and rank it with no inherited votes. Prior variants continue to be listed while eligible.
Statement of Independence
Benchmarking is conducted with strict independence and objectivity. No compensation is received from any providers for listing or favorable outcomes on Artificial Analysis.