Intelligence
that speaks.

Expressive text-to-speech and voice cloning, through one API.

Ready to listen
Transcript Simba 3.2 · Prepared samples

The train slipped out of the city just after sunrise. Maya closed her book and watched the rooftops give way to open fields. For the first time in months, she had nowhere else she needed to be.

Looking for the Speechify reading app? It lives at speechify.com

Speech that keeps up

Ready for real time.
Built for scale.

Fast from the first byte.

56ms

Median time to first audio byte

p50
56 ms
p90
102 ms

Simba 3.2 · Internally measured on our production US East streaming path · 15 Sep 2026. First byte, not audible latency or a guarantee.

Explore streaming

More speech. Less overhead.

$6/ 1M

Per million characters on Scale

Scale subscription$499 / month
Credit in TTS characters78M / month

One $468 shared monthly credit, equivalent to 78M TTS characters at this rate, not a separate character allowance. Continued usage draws from your top-up balance.

Compare plans
Explore the independent benchmarks
Choose the evidence to inspect

Independent blind preference

A frontier voice,
judged without the label.

1,276Elo in the US-accent provider voices view
95% confidence interval
±18
Checked
20 Sep 2026
Samples
1,350

Artificial Analysis Speech Arena

Elo score and 95% confidence interval

Checked 20 Sep 2026
  • SpeechifyAISimba 3.2
    Elo 1276, 95% confidence interval 1258 to 1294.
    1,276 ±18

US-accent provider voices. The reviewed five-model cohort plus two selected alternatives. Lines show 95% confidence intervals, which can overlap.

Inspect the dated figures
More independent evidence315,000 blind preference votes on Datapoint Audio BenchSimba 3.2 entry; exact revision and settings not disclosed · English customer-support prompts · Published 1 Sep 2026
How to read this benchmark

Artificial Analysis runs blinded pairwise listening tests. Simba 3.2 scored 1,276 Elo with a 95% confidence interval of ±18 in this dated view. Its interval overlaps other models in the cohort, so the result does not establish a guaranteed quality advantage.

The first five rows form the economics comparison cohort. ElevenLabs and Gradium are additional selected reference points. The dated figures are SpeechifyAI's reviewed transcription, not an official AA export. “Provisional” is AA's label, not our judgment.

Production US East

Fast from
the first byte.

56 msmedian time to first byte
p50
56 ms
p90
102 ms
Measured
15 Sep 2026

Streaming response

Time to first audio byte on our production path

0 to 150 ms scale

Measured first-byte latency, not audible latency, end-to-end agent response time, or a cross-vendor benchmark.

What this latency measures

These figures measure our Simba 3.2 production streaming path in US East from request to the first audio byte. Network distance, input length, client buffering, and the boundary a benchmark chooses can all change the result.

Time to first byte and time to first audible sound are different measurements. We publish this as a dated production observation, not a guarantee or an independent claim that every route is fastest.

Same reviewed cohort

Frontier quality
without frontier cost.

68%lower than the next-lowest modeled price in the reviewed US-voice cohort
Simba 3.2
$6.60
Next lowest
$20.80
Unit
1M characters

Reviewed five-model cohort, compared on one cost model

USD per 1M characters, zero-based scale

AA estimate · 20 Sep 2026
  1. SpeechifyAISimba 3.2
    $6.60Artificial Analysis estimates SpeechifyAI Simba 3.2 at $6.60 per million characters.
  2. InworldRealtime TTS-2 · provisional
    $20.80Artificial Analysis estimates Inworld Realtime TTS-2 at $20.80 per million characters.
  3. AlibabaQwen-Audio-3.0-TTS-Plus · provisional
    $27.60Artificial Analysis estimates Alibaba Qwen-Audio-3.0-TTS-Plus at $27.60 per million characters.
  4. CartesiaSonic 3.6
    $49.00Artificial Analysis estimates Cartesia Sonic 3.6 at $49.00 per million characters.
  5. VUI LabsLuna TTS
    $80.00Artificial Analysis estimates VUI Labs Luna TTS at $80.00 per million characters.

AA-normalized estimates for the reviewed five-model US-accent provider voices cohort. Subscription assumptions apply. Not a quote or total monthly bill.

Inspect the dated figures
Why modeled price is not your invoice

Artificial Analysis normalizes API pricing to an estimated per-character measure. For subscriptions, it uses the lowest-cost annual plan with at least 1M characters and assumes 80% usage. This chart compares only the reviewed five-model cohort.

The $6.60 figure is that independent model, not a SpeechifyAI billed rate. Our current subscription fees, included usage, and subsequent usage rates are published separately.

Compare voices. Trust your ears.

Compare two prepared voices

Sample
It was the best of times, it was the worst of times, it was the age of wisdom, it was the age of foolishness, it was the epoch of belief, it was the epoch of incredulity.

Two prepared recordings. Choose the voice you prefer.

Voice A
Waiting for your prompt
Voice B
Waiting for your prompt
The details make the difference

More than saying it.
Meaning it.

Identity. Expression. Language.
Explore what makes a voice feel like a voice.

Simba 3.2 · Emotion control

Same words.
Different feeling.

Choose a delivery. Hear how the same voice changes its tone and rhythm.

Neutral.

Beatrice · Delivery sample

The quick brown fox jumped over the lazy dog.

Ready to listen

Prepared Speechify samples. Choose a sample, then press play.Try your own text
The Simba model family

Research you can hear.
Models you can build on.

Explore the models
Simba 3.2

Expression, in real time.

Our streaming-native English model. Shape emotion, pacing, and delivery without losing the character of the voice.

English · SSML control · Voice cloning

Simba 3.0

A voice beyond one language.

Multilingual synthesis for English, German, Spanish, French, Italian, and Portuguese. One streaming-native model.

6 languages · 7 locales · 24kHz audio

From research to your product

Your next idea.
Now with a voice.

Send text. Receive speech. One streaming API, with control over voice, emotion, and timing.

bash
curl https://api.speechify.ai/v1/audio/stream \
  -H "Authorization: Bearer $SPEECHIFY_API_KEY" \
  -H "Content-Type: application/json" \
  -H "Accept: audio/mpeg" \
  -d '{
    "input": "Your next idea. Now with a voice.",
    "model": "simba-3.2",
    "voice_id": "beatrice_32"
  }' --output speech.mp3
Clear from the first character

Start with an idea.
Scale when it speaks.

Compare all plans

Free

Build and ship on the text-to-speech API. No credit card.

$0/month

Start free
Included every month
500K characters
After that
Pauses until next month
  • Catalog voices, streaming, and SSML
  • Commercial use · community support

Starter

For solo devs and early-stage projects.

$10/month

Start Building
Included every month
1.9M characters
After that
$10 per 1M characters
  • Voice cloning, streaming, SSML
  • Keep going past the balance with a top-up, no hard cap

Scale

For serious production speech traffic.

$499/month

Start Scale
Included every month
78M characters
After that
$6 per 1M characters
  • Our lowest published rate
  • Batch synthesis, the only plan that has it

Beyond the standard plans

Enterprise / Custom

For global volume, compliance, and procurement.

Talk to sales

No card needed to start. Voice cloning on paid plans. Included usage renews monthly; purchased top-ups never expire.

A few good questions

Questions,
answered.

Talk with our team
Which model should I start with?

Use Simba 3.2 for new English speech integrations. For supported languages beyond English, use Simba 3.0. Check each voice's supported models before choosing a pairing.

Explore the model family
Can I use my own voice?

Yes. Voice cloning on paid plans creates a reusable voice from a reference recording. Use a voice you own or have the speaker's permission to clone, and check the voice's supported models before synthesis.

Explore voice cloning
Can I try my own text?

The listening demos use prepared recordings. In the SpeechifyAI platform you can try your own text, explore voices, and get an API key for your application.

Explore the developer platform
How does pricing work?

Text-to-speech is metered by character. Starter is $10 per month, including 1.9M characters, with additional usage at $10 per 1M characters from your top-up balance. Free includes 500K characters each month and pauses when that balance runs out.

Compare all plans
Is this the Speechify reading app?

SpeechifyAI is our speech technology and developer platform. The app that reads documents, articles, and books aloud is a separate product at speechify.com.

Visit the reading app

Built on speech. Open to possibility.

Make something
worth listening to.

Privacy preferences

Choose what we may store on this device. You can change this at any time from the footer.

Strictly necessary

Sign-in, security, load balancing, and remembering your privacy choices. These cannot be switched off.

Always on

Analytics

How the site is used in aggregate - which pages get read, where people get stuck - so we can improve it.

Marketing

Measures which campaigns bring people here, and lets us show relevant ads on other platforms.