Fast from the first byte.
56ms
Median time to first audio byte
Simba 3.2 · Internally measured on our production US East streaming path · 15 Sep 2026. First byte, not audible latency or a guarantee.
Explore streamingExpressive text-to-speech and voice cloning, through one API.
The train slipped out of the city just after sunrise. Maya closed her book and watched the rooftops give way to open fields. For the first time in months, she had nowhere else she needed to be.
Looking for the Speechify reading app? It lives at speechify.com
56ms
Median time to first audio byte
Simba 3.2 · Internally measured on our production US East streaming path · 15 Sep 2026. First byte, not audible latency or a guarantee.
Explore streaming$6/ 1M
Per million characters on Scale
One $468 shared monthly credit, equivalent to 78M TTS characters at this rate, not a separate character allowance. Continued usage draws from your top-up balance.
Compare plansIndependent blind preference
Artificial Analysis Speech Arena
Elo score and 95% confidence interval
US-accent provider voices. The reviewed five-model cohort plus two selected alternatives. Lines show 95% confidence intervals, which can overlap.
Inspect the dated figuresArtificial Analysis runs blinded pairwise listening tests. Simba 3.2 scored 1,276 Elo with a 95% confidence interval of ±18 in this dated view. Its interval overlaps other models in the cohort, so the result does not establish a guaranteed quality advantage.
The first five rows form the economics comparison cohort. ElevenLabs and Gradium are additional selected reference points. The dated figures are SpeechifyAI's reviewed transcription, not an official AA export. “Provisional” is AA's label, not our judgment.
Production US East
Streaming response
Time to first audio byte on our production path
Measured first-byte latency, not audible latency, end-to-end agent response time, or a cross-vendor benchmark.
These figures measure our Simba 3.2 production streaming path in US East from request to the first audio byte. Network distance, input length, client buffering, and the boundary a benchmark chooses can all change the result.
Time to first byte and time to first audible sound are different measurements. We publish this as a dated production observation, not a guarantee or an independent claim that every route is fastest.
Same reviewed cohort
Reviewed five-model cohort, compared on one cost model
USD per 1M characters, zero-based scale
AA-normalized estimates for the reviewed five-model US-accent provider voices cohort. Subscription assumptions apply. Not a quote or total monthly bill.
Inspect the dated figuresArtificial Analysis normalizes API pricing to an estimated per-character measure. For subscriptions, it uses the lowest-cost annual plan with at least 1M characters and assumes 80% usage. This chart compares only the reviewed five-model cohort.
The $6.60 figure is that independent model, not a SpeechifyAI billed rate. Our current subscription fees, included usage, and subsequent usage rates are published separately.
Identity. Expression. Language.
Explore what makes a voice feel like a voice.
Simba 3.2 · Emotion control
Choose a delivery. Hear how the same voice changes its tone and rhythm.
Beatrice · Delivery sample
The quick brown fox jumped over the lazy dog.
Ready to listen
Simba 3.2 · Voice cloning
Start with a consented reference recording. Create a reusable voice for the words that come next.
Explore voice cloning01 / Reference
02 / Clone
Two different recordings. Generated sample: Simba 3.2. Only clone a voice you have permission to use.
Simba 3.0 · Multilingual synthesis
Hear 6 language samples, each with a voice from its own catalog. These are separate samples, not translations of one recording.
Explore language supporten-US
de-DE
es-MX
fr-FR
it-IT
pt-BR
Simba 3.2Our streaming-native English model. Shape emotion, pacing, and delivery without losing the character of the voice.
English · SSML control · Voice cloning
Simba 3.0Multilingual synthesis for English, German, Spanish, French, Italian, and Portuguese. One streaming-native model.
6 languages · 7 locales · 24kHz audio
Send text. Receive speech. One streaming API, with control over voice, emotion, and timing.
curl https://api.speechify.ai/v1/audio/stream \
-H "Authorization: Bearer $SPEECHIFY_API_KEY" \
-H "Content-Type: application/json" \
-H "Accept: audio/mpeg" \
-d '{
"input": "Your next idea. Now with a voice.",
"model": "simba-3.2",
"voice_id": "beatrice_32"
}' --output speech.mp3Build and ship on the text-to-speech API. No credit card.
$0/month
Start freeFor solo devs and early-stage projects.
$10/month
Start BuildingFor teams shipping speech in production.
$99/month
Start ProFor serious production speech traffic.
$499/month
Start ScaleBeyond the standard plans
For global volume, compliance, and procurement.
No card needed to start. Voice cloning on paid plans. Included usage renews monthly; purchased top-ups never expire.
Use Simba 3.2 for new English speech integrations. For supported languages beyond English, use Simba 3.0. Check each voice's supported models before choosing a pairing.
Explore the model familyYes. Voice cloning on paid plans creates a reusable voice from a reference recording. Use a voice you own or have the speaker's permission to clone, and check the voice's supported models before synthesis.
Explore voice cloningThe listening demos use prepared recordings. In the SpeechifyAI platform you can try your own text, explore voices, and get an API key for your application.
Explore the developer platformText-to-speech is metered by character. Starter is $10 per month, including 1.9M characters, with additional usage at $10 per 1M characters from your top-up balance. Free includes 500K characters each month and pauses when that balance runs out.
Compare all plansSpeechifyAI is our speech technology and developer platform. The app that reads documents, articles, and books aloud is a separate product at speechify.com.
Visit the reading app