Get started
Published 07.08.2026

How Much Does Realtime TTS Cost? Per-Character, Per-Minute, and Per-User Math

TL;DR: Realtime TTS cost begins with characters or audio tokens, but product economics depend on minutes generated, session length, retries, abandoned audio, pricing tier, and user behavior. Inworld lists TTS-2 from $25 to $5 per million characters and TTS-2 Flash from $15 to $7 before custom enterprise pricing.
A rate per million characters is easy to compare and easy to misuse. Consumer applications pay for repeated interactions, not a static benchmark. The useful calculation converts text into generated minutes, then into sessions, monthly active users, retries, and gross margin. It also separates speech cost from STT, language-model, routing, network, and engineering expense.

How is TTS usage normally priced?

TTS providers commonly bill by characters, credits, or audio-output tokens. Character pricing is the easiest to normalize, but it does not capture speaking rate, failed generations, or unused audio. Teams should record the provider unit, convert it into accepted audio minutes, and attach the contracted pricing tier.
Rates reflect Inworld's public TTS page on September 2, 2026 and can change. ElevenLabs, Cartesia, Deepgram, OpenAI, Google Cloud TTS, Microsoft Azure Speech, and Amazon Polly use different plans and assumptions. Compare the rate actually available to the workload, not an enterprise minimum the team has not qualified for.

How do characters convert into audio minutes?

Characters convert to minutes through speaking rate and language. A practical estimate starts with average characters per spoken minute, then replaces that estimate with measured output from the product corpus. Punctuation, abbreviations, numbers, pauses, and non-verbal cues all change generated duration.
Use the formula: monthly TTS cost equals generated characters divided by one million, multiplied by the contracted rate. Cost per active user equals monthly TTS cost divided by monthly active users. Cost per useful session should exclude tests and failed outputs from the denominator while keeping their expense in the numerator.
  • Measure characters and audio duration together.
  • Separate generated audio from accepted audio.
  • Track retries and cancelled playback.
  • Calculate cost by language and model.
  • Update estimates with production distributions.

What costs sit outside the TTS rate?

TTS is one component of a realtime voice turn. Speech recognition, endpointing, language-model inference, routing, network transport, audio storage, observability, safety review, and engineering labor can exceed the synthesis charge. A reliable business case lists every component and prevents one low vendor rate from hiding a costly system.
Voice agents and companions also generate unused audio when users interrupt. Long buffers can increase waste even when they reduce playback risk. Measure the proportion of synthesized audio that users never hear. The lowest-priced model by character may become expensive when it generates more retries or longer abandoned output.

How do usage tiers change economics?

Volume pricing lowers the unit rate as committed spend or usage grows, but commitments introduce utilization risk. Compare on-demand cost with the effective cost of a plan after unused credits, traffic volatility, and contractual minimums. The best tier is the one with the lowest expected cost under realistic demand, not peak demand.
  1. Model low, expected, and peak usage.
  2. Apply each plan's actual commitment.
  3. Add unused credits as waste.
  4. Include overflow and concurrency charges.
  5. Compare the effective rate per accepted minute.
  6. Recalculate when traffic or pricing changes.
Inworld says selected customers reported cost reductions after changing infrastructure: Talkpal reported 40% lower TTS cost, Bible Chat about 85%, and Wishroll Status about 95% lower AI cost. These are customer-specific outcomes with different workloads and baselines. They should inform test design, not serve as guaranteed savings.

Which product metric should teams optimize?

Consumer applications should optimize cost per useful interaction, completed task, or retained user. Per-character cost is an infrastructure input; it does not show whether the speech improved engagement, reduced abandonment, or supported a viable subscription. Tie model selection to a quality floor and a business outcome.
Talkpal's four-week A/B test linked a speech change to 7% higher feature use and 4% higher retention alongside lower TTS cost, according to Inworld's case study. The strongest experiment randomizes traffic, keeps product conditions stable, and measures cost, quality, latency, completion, and retention together. A lower-priced voice that users abandon is not economically efficient.

Related Guides

Key Takeaways

  • Character prices must be translated into accepted minutes, sessions, and retained users.
  • Retries, failures, interruptions, and unused commitments belong in delivered cost.
  • TTS should be modeled beside STT, LLM, routing, infrastructure, and engineering expense.
  • Customer savings are useful evidence, not guaranteed outcomes for another workload.
  • The strongest decision metric is cost per useful interaction at an acceptable quality floor.

Frequently Asked Questions

How much does Inworld TTS cost?

As of September 2, 2026, Inworld lists TTS-2 at $25 per million characters on demand, $12.50 on the Growth tier, and as low as $5 at enterprise scale. TTS-2 Flash starts at $15 and falls to $7 before custom enterprise pricing. Current rates should be confirmed before purchase.

How many minutes are in one million characters?

There is no universal conversion because speaking rate, punctuation, language, pauses, and non-verbal sounds change duration. Estimate with the intended corpus, then measure characters and generated audio seconds in production. Use the measured distribution rather than a generic average when forecasting monthly spend or cost per user.

Why is cost per character insufficient?

It excludes failures, retries, abandoned audio, unused commitments, speech recognition, language-model inference, routing, networking, and engineering. It also ignores whether users complete the task or return. Cost per accepted minute, useful session, or retained user connects the infrastructure charge to a product outcome.

Do volume plans always reduce cost?

No. A lower listed rate can be offset by unused credits, minimum commitments, traffic volatility, overflow charges, or migration work. Model multiple demand scenarios and calculate the effective rate after waste. On-demand pricing can be cheaper for an uncertain workload even when its nominal unit rate is higher.

Published by Inworld. Prices and product details were checked September 2, 2026 and may change. Customer results are attributed to Inworld-published case studies and should not be generalized without workload-specific testing.
Copyright © 2021-2026 Inworld AI