Intelligence
that speaks.

Expressive text-to-speech and voice cloning, through one API.

Ready to listen
Transcript Simba 3.2 · Prepared samples

The train slipped out of the city just after sunrise. Maya closed her book and watched the rooftops give way to open fields. For the first time in months, she had nowhere else she needed to be.

Looking for the Speechify reading app? It lives at speechify.com

Speech that keeps up

Ready for real time.
Built for scale.

Fast from the first byte.

56ms

Median time to first audio byte

p50
56 ms
p90
102 ms

Simba 3.2 · Internally measured on our production US East streaming path · 15 Sep 2026. First byte, not audible latency or a guarantee.

Explore streaming

More speech. Less overhead.

$6/ 1M

Per million characters on Scale

Scale subscription$499 / month
Credit in TTS characters78M / month

One $468 shared monthly credit, equivalent to 78M TTS characters at this rate, not a separate character allowance. Continued usage draws from your top-up balance.

Compare plans
Explore the independent benchmarks
Choose the evidence to inspect

Independent blind preference

A frontier voice,
judged without the label.

1,237Elo across all provider voices, top five of 92
95% confidence interval
±14
Checked
23 Sep 2026
Board
Top five of 92

Artificial Analysis Speech Arena

Elo score and 95% confidence interval

Checked 23 Sep 2026
  • CartesiaSonic 3.6
    Elo 1273, 95% confidence interval 1256 to 1290.
    1,273 ±17
  • GoogleGemini 3.8 Flash TTS
    Elo 1260, 95% confidence interval 1243 to 1277.
    1,260 ±17
  • AlibabaQwen-Audio-3.0-TTS-Plus
    Elo 1259, 95% confidence interval 1242 to 1276.
    1,259 ±17
  • InworldRealtime TTS-2
    Elo 1245, 95% confidence interval 1227 to 1263.
    1,245 ±18
  • SpeechifyAISimba 3.2
    Elo 1237, 95% confidence interval 1223 to 1251.
    1,237 ±14
  • GoogleGemini 3.8 Flash-Lite TTS
    Elo 1235, 95% confidence interval 1219 to 1251.
    1,235 ±16

The top six of 92 provider voices. Bars show each model's 95% confidence interval; Simba 3.2's overlaps the models around it.

Inspect the dated figures
More independent evidence315,000 blind preference votes on Datapoint Audio BenchSimba 3.2 entry; exact revision and settings not disclosed · English customer-support prompts · Published 1 Sep 2026
How to read this benchmark

Artificial Analysis runs blinded pairwise listening tests. Simba 3.2 scored 1,237 Elo with a 95% confidence interval of ±14 on the dated board, in the top five of 92. Its interval overlaps the models around it, so the result does not establish a guaranteed quality advantage.

The economics view compares the board's top five by Elo. The dated figures are SpeechifyAI's reviewed transcription, not an official AA export.

Production US East

Fast from
the first byte.

56 msmedian time to first byte
p50
56 ms
p90
102 ms
Measured
15 Sep 2026

Streaming response

Time to first audio byte on our production path

0 to 150 ms scale

Measured first-byte latency, not audible latency, end-to-end agent response time, or a cross-vendor benchmark.

What this latency measures

These figures measure our Simba 3.2 production streaming path in US East from request to the first audio byte. Network distance, input length, client buffering, and the boundary a benchmark chooses can all change the result.

Time to first byte and time to first audible sound are different measurements. We publish this as a dated production observation, not a guarantee or an independent claim that every route is fastest.

The board's top five

Frontier quality
without frontier cost.

60%lower than the next-lowest modeled price among the board's top five
Simba 3.2
$6.60
Next lowest
$16.50
Unit
1M characters

The board's top five, compared on one cost model

USD per 1M characters, zero-based scale

AA estimate · 23 Sep 2026
  1. SpeechifyAISimba 3.2
    $6.60Artificial Analysis estimates SpeechifyAI Simba 3.2 at $6.60 per million characters.
  2. GoogleGemini 3.8 Flash TTS
    $16.50Artificial Analysis estimates Google Gemini 3.8 Flash TTS at $16.50 per million characters.
  3. InworldRealtime TTS-2
    $20.80Artificial Analysis estimates Inworld Realtime TTS-2 at $20.80 per million characters.
  4. AlibabaQwen-Audio-3.0-TTS-Plus
    $27.60Artificial Analysis estimates Alibaba Qwen-Audio-3.0-TTS-Plus at $27.60 per million characters.
  5. CartesiaSonic 3.6
    $49.00Artificial Analysis estimates Cartesia Sonic 3.6 at $49.00 per million characters.

AA-normalized estimates for the top five of the 23 Sep 2026 board. Subscription assumptions apply. Not a quote or total monthly bill.

Inspect the dated figures
Why modeled price is not your invoice

Artificial Analysis normalizes API pricing to an estimated per-character measure. For subscriptions, it uses the lowest-cost annual plan with at least 1M characters and assumes 80% usage. This chart compares only the board's top five by Elo.

The $6.60 figure is that independent model, not a SpeechifyAI billed rate. Our current subscription fees, included usage, and subsequent usage rates are published separately.

Compare voices. Trust your ears.

Compare two prepared voices

Sample
It was the best of times, it was the worst of times, it was the age of wisdom, it was the age of foolishness, it was the epoch of belief, it was the epoch of incredulity.

Two prepared recordings. Choose the voice you prefer.

Voice A
Waiting for your prompt
Voice B
Waiting for your prompt
The details make the difference

More than saying it.
Meaning it.

Identity. Expression. Language.
Explore what makes a voice feel like a voice.

Simba 3.2 · Emotion control

Same words.
Different feeling.

Choose a delivery. Hear how the same voice changes its tone and rhythm.

Neutral.

Beatrice · Delivery sample

The quick brown fox jumped over the lazy dog.

Ready to listen

Prepared Speechify samples. Choose a sample, then press play.Try your own text
The Simba model family

Research you can hear.
Models you can build on.

Explore the models
Simba 3.2

Expression, in real time.

Our streaming-native English model. Shape emotion, pacing, and delivery without losing the character of the voice.

English · SSML control · Voice cloning

Simba 3.0

A voice beyond one language.

Multilingual synthesis for English, German, Spanish, French, Italian, and Portuguese. One streaming-native model.

6 languages · 7 locales · 24kHz audio

From research to your product

Your next idea.
Now with a voice.

Send text. Receive speech. One streaming API, with control over voice, emotion, and timing.

bash
curl https://api.speechify.ai/v1/audio/stream \
  -H "Authorization: Bearer $SPEECHIFY_API_KEY" \
  -H "Content-Type: application/json" \
  -H "Accept: audio/mpeg" \
  -d '{
    "input": "Your next idea. Now with a voice.",
    "model": "simba-3.2",
    "voice_id": "beatrice_32"
  }' --output speech.mp3
Clear from the first character

Start with an idea.
Scale when it speaks.

Compare all plans

Free

Build and ship on the text-to-speech API. No credit card.

$0/month

Start free
Included every month
500K characters
After that
Pauses until next month
  • Catalog voices, streaming, and SSML
  • Commercial use · community support

Starter

For solo devs and early-stage projects.

$10/month

Start Building
Included every month
1.9M characters
After that
$10 per 1M characters
  • Voice cloning, streaming, SSML
  • Keep going past the balance with a top-up, no hard cap

Scale

For serious production speech traffic.

$499/month

Start Scale
Included every month
78M characters
After that
$6 per 1M characters
  • Our lowest published rate
  • Batch synthesis, the only plan that has it

Beyond the standard plans

Enterprise / Custom

For global volume, compliance, and procurement.

Talk to sales

No card needed to start. Voice cloning on paid plans. Included usage renews monthly; purchased top-ups never expire.

From the blog

Start here:
choosing a TTS API.

Read the blog
  1. How good does it sound?

    1,237Elo

    Best TTS APIs 2026: 10 models compared by Elo and price

    16 min read, 15 min listen

  2. What does it cost?

    $91a month for 10M on Starter

    TTS API pricing 2026: what 1M, 10M and 100M characters a month really cost

    7 min read, 5 min listen

  3. How fast does it start?

    123ms to first audible audio

    How to choose a low-latency TTS API for voice agents in 2026

    10 min read, 8 min listen

A few good questions

Questions,
answered.

Talk with our team
Which model should I start with?

Use Simba 3.2 for new English speech integrations. For supported languages beyond English, use Simba 3.0. Check each voice's supported models before choosing a pairing.

Explore the model family
Can I use my own voice?

Yes. Voice cloning on paid plans creates a reusable voice from a reference recording. Use a voice you own or have the speaker's permission to clone, and check the voice's supported models before synthesis.

Explore voice cloning
Can I try my own text?

The listening demos use prepared recordings. In the SpeechifyAI platform you can try your own text, explore voices, and get an API key for your application.

Explore the developer platform
How does pricing work?

Text-to-speech is metered by character. Starter is $10 per month, including 1.9M characters, with additional usage at $10 per 1M characters from your top-up balance. Free includes 500K characters each month and pauses when that balance runs out.

Compare all plans
Is this the Speechify reading app?

SpeechifyAI is our speech technology and developer platform. The app that reads documents, articles, and books aloud is a separate product at speechify.com.

Visit the reading app

Built on speech. Open to possibility.

Make something
worth listening to.

Privacy preferences

Choose what we may store on this device. You can change this at any time from the footer.

Strictly necessary

Sign-in, security, load balancing, and remembering your privacy choices. These cannot be switched off.

Always on

Analytics

How the site is used in aggregate - which pages get read, where people get stuck - so we can improve it.

Marketing

Measures which campaigns bring people here, and lets us show relevant ads on other platforms.