ElevenLabs: What It Is, Pricing and How to Use It (2026 Guide)
elevenlabs.ioElevenLabs generates realistic AI voices for narration, dubbing and voice agents. What it does, what it costs, and a step-by-step guide to your first voiceover.
PromptWises is reader-supported. If you sign up through links on this page we may earn a commission, at no extra cost to you. This does not influence what we write.
What is ElevenLabs?
ElevenLabs is an AI audio platform best known for text to speech that sounds like a person reading rather than a computer. You type or paste a script, pick a voice, and it produces a narration you can download as MP3 or WAV. Around that core it has built a wide set of audio tools: voice cloning, voice design from a text description, dubbing of video into other languages, speech to text, music and sound effect generation, and a platform for building conversational voice agents.
The company runs everything in the browser at elevenlabs.io, with an API and SDKs for developers. The same account and the same pool of credits cover every product, so a creator can narrate a video, dub it and transcribe it without buying three subscriptions.
The problem it solves is simple. Good voiceover used to mean a voice actor, a studio and days of turnaround, or a robotic text-to-speech engine nobody wanted to listen to. ElevenLabs gets close enough to a human read that most listeners do not notice, and it does it in seconds.
Who it is for
- YouTubers, course creators and explainer-video producers who narrate a script every week and want consistent, editable voiceover without recording.
- Podcasters and audiobook producers who need long-form narration with chapters and more than one speaker.
- Marketing and localisation teams that dub product videos and ads into other languages while keeping the original speaker’s voice.
- Developers building voice agents, reading apps or interactive characters, where low latency and an API matter more than the web editor.
- Accessibility and internal teams turning documents and training content into audio.
ElevenLabs is not the right choice if you only need a quick, cheap voice for an internal video and do not care about naturalness; simpler tools cost less. It is also not a video editor. You will export audio and finish the project elsewhere.
Key features
- Text to Speech. Paste up to 5,000 characters in the web app (2,500 on the free plan), choose a voice and a model, and generate. The models differ in trade-offs: Eleven v4 and Eleven v3 are the most expressive, Eleven Multilingual v2 is the most stable for long reads in 29 languages, and Eleven Flash v2.5 is a low-latency model for real-time use.
- Audio tags. With the v4 and v3 models you can put short directions in square brackets inside the script, such as
[whispers],[laughs]or[sighs], and the voice follows them. - Voice Library. More than 10,000 community-shared voices you can search by language, accent, age and style, preview, and add to My Voices.
- Voice Design. Describe a voice in 20 to 1,000 characters (age, gender, accent, tone) and ElevenLabs generates three previews to pick from.
- Voice cloning. Instant Voice Cloning works from one to two minutes of clean audio and is available from the Starter plan. Professional Voice Cloning uses 30 to 180 minutes of recordings, is available from Creator, and produces a higher-fidelity clone after a few hours of training.
- Studio. A long-form editor for audiobooks, podcasts and narrated videos. Import an EPUB, PDF, DOCX, TXT, HTML file or URL, split it into chapters, assign different voices to different paragraphs, and export per chapter or as one file.
- Dubbing. Upload a video or paste a YouTube or TikTok link and get a translated, re-voiced version in 90 or more languages, with the original speakers’ voices preserved.
- Speech to Text, Music and Sound Effects. Scribe transcribes in 90 or more languages with speaker labels; Eleven Music generates tracks from a text prompt; Sound Effects generates short effects from a description.
- Agents and API. ElevenAgents builds conversational voice agents for phone and web; ElevenAPI exposes everything above to developers with Python and TypeScript SDKs.
Pricing
ElevenLabs sells a monthly credit allowance. Text to speech costs one credit per character on the standard models, so 10,000 credits is roughly ten minutes of speech. The prices below are the regular monthly prices at the time of writing; yearly billing means you pay for ten months and get twelve. Check the pricing page before buying, as ElevenLabs runs frequent first-month discounts.
| Plan | Monthly price | Credits per month | Custom voice slots | Voice cloning | Notes |
|---|---|---|---|---|---|
| Free | $0 | 10,000 (about 10 minutes) | 3 | None | No commercial licence, watermarked dubs, 5 Studio projects |
| Starter | $6 ($5 billed yearly) | 30,000 (about 30 minutes) | 10 | Instant | Commercial licence, 20 Studio projects |
| Creator | $22 ($18.33 billed yearly) | 121,000 (about 2 hours) | 30 | Instant plus 1 Professional | 192 kbps audio, 1,000 Studio projects |
| Pro | $99 ($82.50 billed yearly) | 600,000 (about 10 hours) | 160 | Instant plus 1 Professional | 44.1 kHz PCM output via API |
| Scale | $299 | 1.8 million | 660 | Instant plus 3 Professional | 3 seats |
| Business | $990 | 6 million | 2,200 | Instant plus 10 Professional | 10 seats, low-latency TTS from about 5 cents a minute |
| Enterprise | Custom | Custom | Custom | Custom | SSO, DPA and SLA terms, HIPAA BAA, managed dubbing |
Other products draw on the same credits at different rates: speech to text costs 330 credits per minute, music 900 per minute, sound effects 200 per generation, voice changer and voice isolator 1,000 per minute, and dubbing between 2,000 and 10,000 per minute depending on whether you use automatic dubbing or Dubbing Studio and whether you accept a watermark. Unused credits on paid plans roll over for up to two months, capped at twice the monthly quota. Extra credits can be bought as pay-as-you-go top-ups.
How to use ElevenLabs: step-by-step
1. Create an account and pick a voice
- Go to elevenlabs.io and sign up with Google or email. The free plan needs no card.
- Open Voices in the left sidebar. The Voice Library tab lists community voices; filter by language, gender, age and use case, press play on a few, and click Add to Collection on the ones you like. They appear under My Voices.
- Voice choice has the largest effect on the result. Test three or four voices on the same two sentences before you settle.
2. Generate your first narration
- Click Text to Speech in the sidebar.
- Paste your script into the text box. Write out numbers, currencies and abbreviations as words (“one hundred dollars”, not “$100”) for reliable pronunciation.
- Select a voice from your voices at the bottom left of the screen.
- Choose a model from the model selector. For a YouTube narration or course module, start with Eleven v4 or Eleven Multilingual v2. For a short, dramatic read, try Eleven v3.
- Leave the voice settings at their defaults for the first run. The main ones are Stability (lower gives more emotional range, higher gives a more consistent but flatter read, around 50 is a sensible start), Similarity, Style exaggeration (keep at 0 unless you have a reason), Speed (0.7 to 1.2) and Speaker Boost.
- Click Generate. The audio appears in a player below the text.
- If the text, voice and model are unchanged you usually get two free regenerations, so adjust Stability or Speed and click Regenerate speech if the first take is off.
- Click the download button at the bottom right for MP3 at 128 kbps or WAV. Click Advanced for 192 or 256 kbps MP3, M4A or FLAC. Every generation is also saved in the history panel on the right, where you can download it later.
3. Direct the delivery with audio tags
With Eleven v4 or Eleven v3 selected, put stage directions in square brackets where you want the delivery to change:
[excited] We just crossed ten thousand subscribers.
[pause] Thank you.
[whispers] And there is one more thing I have not told you yet.
Use tags sparingly. One tag every few sentences does more than a tag on every line, and some tags work better with some voices than others. Generate, listen, and remove tags that make the read worse.
4. Clone your own voice
- Record one to two minutes of yourself speaking naturally in a quiet room. More than three minutes adds little. Use the same tone you want the clone to produce.
- Open Voices, click the plus icon, and choose Instant Voice Clone.
- Upload or record the audio, give the voice a name, tick the box confirming you have the right to clone this voice, and click Save voice.
- The clone appears under My Voices. Click Use voice and generate a test paragraph. If it sounds thin, re-record with a better microphone rather than fiddling with settings.
Professional Voice Cloning (Creator plan and above) follows a similar path but asks for 30 minutes to three hours of audio, verifies that the voice is yours, and emails you when training is done, typically within a few hours.
5. Build a long-form project in Studio
- Open Studio and either describe the project in the prompt bar, click Upload to import a script (EPUB works best because chapters are detected from Heading 1 styles), or choose + New blank project and then Audio project.
- Use the Chapters sidebar to add, rename and reorder chapters.
- Select any paragraph and pick a voice from the drop-down to assign a different speaker. Use Override settings to change voice settings for a selection only, and Replace voice across project to swap a voice everywhere.
- Click Export to render a chapter or the whole project as MP3 or WAV. Multi-chapter projects can be exported as a ZIP with one file per chapter. Free, Starter and Creator plans export at 128 kbps; Pro and above get 192 kbps MP3 or 16-bit 44.1 kHz WAV.
6. Dub a video into another language
- Open Dubbing from the navigation menu.
- Upload a video or audio file (up to 2 GB and 180 minutes on the website) or use the Paste URL tab for a YouTube or TikTok link.
- Pick target languages under Choose languages.
- Adjust Speaker similarity if you want the dubbed voice closer to or further from the original; the default of 7 suits most content.
- Click generate, review the credit cost shown, and confirm. Download the finished dub from your list of dubs when it is ready. Free-plan dubs carry a watermark that cannot be removed.
Tips to get better results
- Spend credits on the voice choice, not on settings. Ten short test generations across different voices will teach you more than an hour of slider adjustments.
- Write for the ear. Short sentences, contractions and spelled-out numbers produce more natural reads. Read your script aloud once before generating.
- Use punctuation to control pacing. Commas, full stops and line breaks create pauses. An ellipsis creates a longer one.
- Keep Stability in the middle for narration. Very low values sound lively but can drift between takes; very high values sound flat.
- Match the model to the job. Multilingual v2 for long, even reads; v4 or v3 for emotion; Flash v2.5 or v4 Turbo for anything real-time.
- Regenerate the sentence, not the paragraph. In Studio you can select one paragraph and regenerate it, which costs a fraction of re-rendering the chapter.
- Fix pronunciation at the source. If a brand name is mispronounced, respell it phonetically in the script.
- Watch the credit counter. Dubbing and music are far more expensive per minute than plain speech. Check the cost preview before confirming a dub.
Limitations and common problems
- Credits go quickly on small plans. A 1,500-word script is roughly 9,000 characters, which is almost the whole free allowance. Budget for Starter or Creator if you publish weekly.
- Occasional mispronunciations and odd emphasis. Even the best voices stumble on unusual names or long numerical strings. Respell the word or split the sentence and regenerate.
- Consistency across long reads. Expressive models can shift tone between paragraphs. For audiobooks, Multilingual v2 at a higher Stability is more even.
- Cloning quality depends on the sample. Background noise, compression and room echo transfer into the clone. Record close to a decent microphone in a quiet room.
- Dubbing is not a finished master. The dubbed video is a preview; for production, download the audio and replace the track on your original video in an editor. Automatic dubbing has no in-app transcript editing.
- Consent and licensing. You must have the right to clone any voice you upload, and free-plan output has no commercial licence. Read the terms before using a clone of anyone but yourself.
- Browser-only editing. There is no offline desktop app for the editor, so download your files before you need them somewhere without a connection.
Alternatives
- Murf AI - a studio-style editor aimed at corporate video, e-learning and ads, with timeline-based media and simpler per-minute pricing. See our Murf AI guide.
- OpenAI text to speech and Google Cloud Text-to-Speech - cheaper API options for developers who need volume over expressiveness.
- Descript - better if you want to edit a video or podcast by editing the transcript, with its own voice cloning built in.
- Play.ht and WellSaid Labs - established text-to-speech tools with large voice catalogues, popular with e-learning teams.
- Adobe Podcast and Audacity - if you are recording your own voice, clean-up tools may serve you better than synthesis.
Browse more options in our AI content and media tools directory. For scripting help, our guide on how to write better prompts covers how to get a usable first draft from ChatGPT or Claude before you hand it to a voice.
Verdict
ElevenLabs is the right choice when the voice itself matters: narration people will listen to for minutes at a time, dubbing where the original speaker’s identity should survive translation, or agents that need to sound human. The free plan is enough to judge quality on your own script in an afternoon. Creators publishing regularly should expect to need Creator or above; occasional users can stay on Starter. If you mainly need voiceover laid over slides and stock video in a single editor, look at Murf AI first.
Frequently asked questions
Is ElevenLabs free?
Yes. The free plan gives you 10,000 credits a month, which is roughly ten minutes of text to speech, plus three custom voice slots. It does not include a commercial licence or voice cloning, and dubs carry a watermark. Paid plans start with Starter at $5 a month billed yearly at the time of writing.
Can I use ElevenLabs voices commercially?
A commercial licence is included from the Starter plan upward. Free-plan output is for non-commercial use. Check the current terms on the pricing page before publishing paid work.
How does ElevenLabs voice cloning work?
Instant Voice Cloning needs one to two minutes of clean audio and is available from the Starter plan. Professional Voice Cloning uses 30 to 180 minutes of recordings, requires the Creator plan or higher, includes voice verification, and takes a few hours to train.
What are ElevenLabs credits?
Credits are the usage unit shared across all products. Text to speech costs one credit per character on the standard models. Dubbing, transcription, music and sound effects cost a set number of credits per minute or per generation. Unused credits on paid plans roll over for up to two months.
Which ElevenLabs model should I use?
Use Eleven v4 or Eleven v3 for expressive narration where you want control over delivery, Eleven Multilingual v2 for long, consistent reads, and Eleven Flash v2.5 or v4 Turbo when latency matters, such as in voice agents.
Does ElevenLabs do dubbing?
Yes. Upload a video or paste a YouTube link, choose target languages, and ElevenLabs translates and re-voices it while keeping the original speakers' voices. It supports 90 or more languages and costs between 2,000 and 10,000 credits per minute depending on the options.