Voice identity · API

Your voice.
A new possibility.

Create a synthetic voice from a 10-30 second clip and verified spoken consent, then synthesize text in that voice by ID.

Account setup and your API key are free. Voice cloning requires Starter or above.

Create a voiceConsent required
bash
curl -X POST https://api.speechify.ai/v1/voices \
  -H "Authorization: Bearer $SPEECHIFY_API_KEY" \
  -H "Speechify-Version: 2026-09-13" \
  -F name="Narrator" \
  -F gender="female" \
  -F sample=@sample.wav \
  -F consent_challenge_id="$CONSENT_CHALLENGE_ID" \
  -F consent_recording=@consent.wav

First create a consent challenge with full_name, then record the speaker reading its phrase. Set CONSENT_CHALLENGE_ID to the returned ID and save the recording as consent.wav before submitting the sample.

  • 10-30s voice sample
  • Instant, self-serve
  • 6 self-serve languages on Simba 3.0
  • Zero-shot and fine-tuned
Voice cloning · Listening samples

Hear a clone before you write a line.

Listen to the original speaker and a prepared cloned-voice sample generated with Simba 3.0. Two recordings, different scripts, the same voice identity.

Simba 3.0 · Voice cloning

A new script.
Your voice.

Start with a consented reference recording. Create a reusable voice for the words that come next.

Explore voice cloning

01 / Reference

Original speaker

Ready to listen0:00

02 / Clone

Cloned voice

Ready to listen0:00

Two different recordings. Generated sample: Simba 3.0. Only clone a voice you have permission to use.

Prepared Speechify samples. Choose a sample, then press play.Try your own text
How it works

Three steps to a cloned voice.

No separate service to integrate. Cloning is an endpoint alongside synthesis on the Build API.

  1. Record a sample

    Prepare 10-30 seconds of clean speech from the person whose voice you are authorized to clone.

  2. Verify consent

    POST to /v1/voices/consent-challenges with the speaker's full name as the JSON field full_name, then record the same speaker reading the returned phrase.

  3. Create and synthesize

    Submit the sample, consent recording, and challenge ID to /v1/voices. Use the returned voice ID with a supported model on the speech endpoints.

Quality tiers

Two tiers, one API.

Start self-serve, or work with our team on a fine-tuned voice.

Instant · zero-shot

Self-serve cloning

Clone from a 10-30 second sample and a verified consent recording via the API or Console. No fine-tuning run required.

Instant cloning
Fine-tuned

Professional cloning

Fine-tune on hours of a speaker's audio. Arranged with our team for signature narrators and brand voices.

Arrange professional cloning

Both Simba 3.2 and Simba 3.0 are streaming-native and support self-serve voice cloning. Use Simba 3.2 for English, or Simba 3.0 across 6 languages and 7 locales. Check each voice's supported models before synthesis; discuss broader language requirements with our team.

Pricing

Part of Build, on one bill.

Voice cloning is included on paid plans, with no separate per-clone charge. Synthesis uses your plan's per-character rate, just like a catalog voice. Free accounts do not include cloning.

cloning included
Starter
per 1M chars · synthesis
from $6
voice sample guidance
10-30s
Simba 3.0 self-serve languages
6
Use cases

Where a cloned voice earns its keep.

Accessibility

Assistive technology that speaks in a person's own voice instead of a generic one: clone from a short sample with consent and use the voice ID in read-aloud and communication tools, or preserve a voice before it is lost.

Audiobooks

An author or narrator reads a whole book in their own voice, generated from the manuscript. One voice ID keeps every chapter consistent, and the same clone reads the translated editions.

Creator tools

A cloning feature inside your own product: your users clone their voice and generate content in it. The API provides the endpoints and enforces consent; you provide the flow and manage the voice IDs.

Podcasts

One signature host voice across every episode, produced from a script, with cross-language editions in the same voice rather than a different presenter per market.

Game characters

A signature character keeps one consistent voice across every line and every update, with new lines synthesized on demand from a consented clone of the actor.

FAQ

Common questions.

What is voice cloning?
It creates a synthetic version of a specific voice from an audio sample, then synthesizes any text in that voice. On the Build API you clone from a short clean sample, with the speaker's consent, and use it by voice ID on the normal speech endpoints, a Build feature that shares the Build API and pricing.
Do I need consent to clone a voice?
Yes. Create a consent challenge and record the speaker reading its phrase. Submit that recording with the challenge ID and voice sample. The recording is checked against the phrase and speaker, then retained as the consent record.
What languages do cloned voices support?
Simba 3.0 supports 6 languages across 7 locales self-serve. Simba 3.2 is English only. Check each voice's supported models before synthesis and discuss broader language requirements with our team.
Instant or professional cloning: which one?
Instant cloning is zero-shot: a 10-30 second voice sample and a verified consent recording create a reusable voice without a fine-tuning run. Professional cloning fine-tunes on hours of a speaker's audio and is arranged with our team for signature narrators and brand voices.
What does voice cloning cost?
Nothing separate. Cloning is a Build capability included on Starter, Pro, Scale and Enterprise (a Free-tier attempt returns a 402 with the code voice_cloning_not_included), and speaking with a cloned voice is billed per character like any catalog voice, from $6 per 1M characters. There is no per-clone price and no synthesis premium.
How do I manage cloned voices?
Retrieve a voice with GET /v1/voices/{voice_id}, download its sample as WAV, and remove it with DELETE /v1/voices/{voice_id}. Cloned voices list before shared voices.

Build with a verified voice.

Create an account and API key, choose a paid plan, then record a sample and verify the speaker's consent.

Account setup is free. Voice cloning requires Starter or above.

Privacy preferences

Choose what we may store on this device. You can change this at any time from the footer.

Strictly necessary

Sign-in, security, load balancing, and remembering your privacy choices. These cannot be switched off.

Always on

Analytics

How the site is used in aggregate - which pages get read, where people get stuck - so we can improve it.

Marketing

Measures which campaigns bring people here, and lets us show relevant ads on other platforms.