ScruTool

ElevenLabs AI

ElevenLabs is an AI voice platform with the most convincing synthetic speech available, plus music, dubbing and video. Everything runs off one credit meter at very different rates.

Arden Presley
Reviewed by
Arden Presley · Tech Reviewer
Updated On
Aug 20, 2026
Scrutool Score
6.7 /10
★★★☆☆
Worth a try

The most convincing synthetic voice available, sold through a credit meter that makes experimenting expensive and downgrading painful.

Voice realism & expressiveness 9.0
Emotional control via audio tags 8.5
Breadth of creative tools 8.0
Output consistency across regenerations 5.5
Free tier practicality 5.0
Billing transparency & support 4.5
Category AI Voice Generator Text to Speech Voice Cloning AI Music Dubbing
Platform Type
Freemium web · API
Free Credits
10k/month about 10 min speech
Commercial License
Paid Starter and up
Music Download
Paid free is preview only
Video Generation
Paid Starter and up
Credit Rollover
Up to 2 months paid plans only
30-Second Verdict

Should you use ElevenLabs?

Best for

Voiceover, audiobooks and character work where delivery matters. Nothing else currently renders a whisper or a shout this convincingly.

Skip if

Your workflow needs identical audio on every regeneration . The same text, voice and model returned visibly different waveforms in testing.

Real cost

Free gives 10,000 credits and no commercial license. Exporting music or touching video starts at Starter, $6/mo .

Watch out

Credits roll over for two months, but downgrading or cancelling forfeits them at the end of the cycle.

Overview

What is ElevenLabs?

One meter runs the whole platform. Speech, music, sound effects, dubbing, images and video all draw from a single credit balance, and they draw at rates that are nowhere near each other. A minute of speech costs about 1,000 credits. Thirty seconds of music costs 450. A second of video, by the platform's own arithmetic on the upgrade screen, costs roughly 172. Understanding that spread is most of what deciding on ElevenLabs involves, because the tool you came for and the tool that empties your balance are usually not the same one.

What the credits buy is the most convincing synthetic voice on the market. ElevenLabs runs on the web and through an API, offers 70+ languages and a library of thousands of voices, and its Eleven v3 model takes bracketed audio tags like [whispers] and [sighs] as performance direction rather than text to read. It suits voiceover, audiobooks, dubbing and character work. Free accounts get 10,000 credits a month and no commercial license.

Capabilities

Key features & how they perform

Each feature rated from hands-on testing and aggregated review sentiment.

🎙️
★★★★★ 4.7

Text to speech realism

The category benchmark. G2 reviewers describe prosody that holds across long scripts, and most complaints on review sites are about cost rather than the voices.

🏷️
★★★★☆ 4.3

Audio tags in Eleven v3

Bracketed cues are read as direction, not dialogue. Reliability improved since launch, though some users still report tags firing inconsistently across voices.

🧬
★★★★☆ 4.4

Voice cloning

Instant cloning from Starter, Professional cloning from Creator. Consent verification is required, and clone quality tracks closely with sample length and recording conditions.

🎵
★★★★☆ 3.9

Eleven Music

Prompt adherence was strong in testing on instrumentation and texture. The cost is quoted before you commit, and free accounts cannot export what they make.

🌍
★★★★☆ 3.5

Multilingual coverage

70+ languages on paper. English is the strongest by a distance, and reviewers repeatedly single out Japanese and regional German variants as noticeably weaker.

🎬
★★★☆☆ 2.8

Image and video generation

Talking avatars and third-party video models sit inside the same workspace, but the credit rate is roughly ten times that of speech and none of it reaches the free plan.

Feature ratings blended from G2, Capterra, Trustpilot and Reddit review patterns plus hands-on testing.

Hands-On Walkthrough

What 10,000 free credits actually bought me

Audio tags, a number test, a music track and two paywalls, all on one free account.

Open a free account and find the balance

ElevenLabs gives free accounts 10,000 credits a month, which the pricing page translates to roughly ten minutes of text to speech.

The balance is not on the main workspace. It lives in the account menu in the top corner, alongside the workspace switcher and subscription link, which is a strange place for the number that governs everything you can do.

Free accounts do not carry a commercial license, and credit rollover does not apply to them.

During testing Signup asked for no card and started no trial countdown. The account opened on 10,000 credits with Multilingual v2 preselected as the model, which matters more than it looks, as the next step shows.

Push the audio tags in Eleven v3

Audio tags are the reason Eleven v3 exists. You write a cue in square brackets and the model treats it as direction for the performance rather than words to speak.

Three cues in one short line is a harder test than it sounds, because the model has to change register twice inside two sentences.

The input

"[whispers] I can't believe you did that. [sighs] It's over. [shouts] Just go!"

Getting to that generation took a detour. My first attempt was still sitting on Multilingual v2, and instead of running it, the platform stopped me.

What surprised me Seventy-seven characters. The whisper came out breathy and dropped in volume, the sigh registered as audible breath rather than the word being read aloud, and the shout carried the strain a real voice produces when it is raised. The detour mattered too: a modal explained that audio tags only work in Eleven v3 and offered one button to switch. Most tools in this category would have silently generated a voice saying "open bracket whispers close bracket" and charged for it.

Re-run the number test everyone still quotes

A widely circulated 2025 review found ElevenLabs reading 200000 as "twenty thousand thousand". That failure still gets repeated in reviews published this year.

I gave v3 both the plain six-digit number and an Indian lakh grouping in one line, since the second is the harder case and almost never gets tested.

Exact prompt

"I have 200000 coins and 1,23,456 rupees."

The result Both were handled. The large number came out as two hundred thousand, and 1,23,456 was read as one lakh twenty-three thousand four hundred fifty-six rather than being flattened into a Western reading. Reviews still repeating that complaint have not retested it. The screenshot shows something else: Generation 1 and Generation 2 came from identical text, voice, model and stability setting, eleven seconds apart, and the waveforms do not match. For a one-off voiceover that is invisible. For a series where you regenerate one corrected line and drop it back into a finished edit, it is a problem to work around.

Generate a music track and watch the price first

Music is included on the free plan, which is more than most competitors offer at zero cost.

The composer quotes the credit cost next to the length control before you commit, so you can see what a track will spend while you are still writing the prompt. I gave it something specific enough to fail at.

What I asked for

"Slow lo-fi hip hop with sitar, rainy afternoon mood, with vinyl crackle and a mellow Rhodes piano melody"

The finished track landed in the project history with a name and genre tags it had written itself.

What came back Thirty seconds was priced at 450 credits before I committed, which is the same transparency the speech side gets right. The track arrived in under a minute, auto-titled "Cobblestone Reflections" and tagged as lo-fi hip hop and ambient. The sitar was there. The vinyl crackle was there. The Rhodes tone sat about where I asked it to.

Try to download what you just paid for

The download icon sits on the player like any other control, with nothing marking it as restricted.

Clicking it is where the free music allowance reveals what it actually is.

Doing the math Downloading music requires a paid subscription. The generation had already spent 450 credits, 4.5 percent of the entire monthly free allowance, on a file that never leaves the browser. A free user can produce about twenty-two tracks a month and export none of them. That reframes the free music tier as a demo surface rather than a working allowance.

Meet the video upsell inside the speech workflow

ElevenLabs now carries image and video generation alongside audio, with third-party models available in the same workspace.

The promotion for it does not sit in a separate tab. It appears on the completion screen for a finished speech clip, directly beside the regenerate button, offering to turn the clip into a talking avatar video.

Clicking it opens a modal that prints the credit arithmetic in full, which turns out to be the most useful thing on the screen.

Worth knowing Video generation is off the free plan entirely. The Creator tier at $22 a month, halved to $11 for a first month, carries 121,000 credits, which the modal itself prices at up to 660 images or 705 seconds of video. Run that division and a second of video costs roughly 172 credits, against the 1,000 credits that buy a full minute of speech. Video is priced at about ten times the rate of the product the platform is known for.
Plans & Cost

ElevenLabs pricing

Figures taken from the official ElevenCreative pricing page and help centre.

Plan Price What's included
Free $0 10k credits/month (~10 min speech) · speech, music, SFX, voice design, image · 3 Studio projects · no commercial license · no rollover
Starter $6 /mo 30k credits (~30 min) · commercial license · music commercial use · instant voice cloning · dubbing studio · image & video · 20 Studio projects
Creator Popular $22 /mo First month $11 · 121k credits (~121 min) · professional voice cloning · pay-as-you-go top-ups · everything in Starter
Pro $99 /mo 600k credits (~600 min) · 192kbps audio · 44.1kHz PCM output via API · everything in Creator
Scale $299 /mo 1.8M credits (~1,800 min) · 3 workspace seats · 3 professional voice clones · team collaboration
Business $990 /mo 6M credits (~6,000 min) · 10 workspace seats · 10 professional voice clones · low-latency TTS from 5c/minute
Enterprise Custom Contact sales · custom DPA/SLA terms · BAAs for HIPAA · custom SSO · managed dubbing with Productions

Annual billing is priced as ten months rather than twelve, which works out to $5, $18.33, $82.50, $249.17 and $825 a month across the five paid tiers.

⚠ Where the bill goes sideways The free plan carries no commercial license, cannot export music, and cannot touch video. Creator's $11 is a first-month rate that reverts to $22. Video burns credits at roughly ten times the rate of speech, so a plan sized for voiceover empties fast the moment you try it. Credits roll over for up to two months on an active paid plan, but downgrading or cancelling forfeits whatever is unused at the end of the cycle. Trustpilot reviewers repeatedly report charges continuing after they believed they had cancelled.
The Balance

Pros & cons

Specific conclusions from testing and real user reviews, not generic filler.

✅ Pros

  • Audio tags render whispers and shouts as performance, not spoken text
  • A modal blocks tag use on the wrong model before spending credits
  • Music cost is quoted upfront: 450 credits for a 30 second track
  • Eleven v3 handled 200000 and the lakh grouping 1,23,456 correctly
  • Free account opens with 10,000 credits and no card requested
  • 70+ languages and thousands of voices in one library
  • Paid credits roll over up to two months on an active plan

⛔ Cons

  • Identical text on the same voice and model returned different waveforms
  • Free accounts cannot download generated music at all
  • One 30 second track costs 4.5% of the monthly free allowance
  • Video runs about 172 credits a second against 1,000 per speech minute
  • Creator's $11 rate covers the first month only, then $22
  • Downgrading or cancelling forfeits unused credits at cycle end
  • Trustpilot sits near 3.2 against 4.5 on G2, driven by billing
  • Reviewers report automated support replies before any escalation
  • Japanese and regional German output lag English noticeably

Synthesized from real reviews on Trustpilot, G2, Capterra and Reddit · paraphrased, not quoted

Benchmarks

ElevenLabs scorecard

Rated against what an AI voice and audio platform is actually built to do.

How we score Each dimension is rated 0 to 10 from hands-on testing combined with aggregated user-review sentiment (G2, Capterra, Trustpilot, Reddit) and the platform's published documentation. The headline Scrutool Score is the equal-weight average of all 10 dimensions below.
Dimension Verdict Score
Voice realism & expressiveness Naturalness, prosody, delivery Excellent
9.0
Emotional control via audio tags Whisper, sigh, shout accuracy in v3 Excellent
8.5
Breadth of creative tools Speech, music, SFX, dubbing, video Good
8.0
Text normalization & pronunciation Numbers, currency, heteronyms Good
7.5
Interface & workflow design Guardrails, cost visibility, navigation Good
7.5
Language coverage & non-English quality Realism outside English Average
6.5
Output consistency across regenerations Same input, same result Average
5.5
Free tier practicality What 10k credits actually delivers Weak
5.0
Credit value & cost predictability Rate spread across products Weak
5.0
Billing transparency & support Cancellations, forfeits, response quality Weak
4.5
Scrutool Score Equal-weight average of all 10 dimensions
6.7
Sentiment Analysis

What users say about ElevenLabs

The themes reviewers raise most often, by share of analysed reviews.

👍 Most-mentioned praise
Voice realism ahead of every alternative tested 79%
Fast setup and a working first generation in minutes 52%
Speech, music, dubbing and video in one platform 44%
Voice cloning quality from short samples 38%
Support resolves generously once a human is reached 26%
👎 Most-mentioned pain
Credits drain faster than the plan suggests 61%
Billing and cancellation problems 47%
Paid credits forfeited on downgrade or cancellation 35%
Output varies between regenerations of the same line 33%
Automated support replies before any human contact 29%

% = share of analysed reviews mentioning each theme (Trustpilot, G2, Capterra, Reddit)

6.7
Final Verdict

The best voice on the market, sold through a meter that punishes experiments

Nothing else does what Eleven v3 does with a bracketed cue. The whisper drops in volume, the sigh arrives as breath, the shout carries actual strain, and a modal stopped me from wasting credits on the wrong model before I knew I had picked it. That quality of engineering is real. The commercial design around it is harder to like. A single 30 second music track costs 4.5 percent of the free monthly allowance and cannot be exported at all. Video runs at roughly ten times the credit rate of speech. Creator advertises $11 and charges $22 from month two. Downgrade or cancel and the credits you already paid for disappear at the end of the cycle, which is the single loudest complaint on Trustpilot and the reason its score sits so far below the same product's rating on G2. Voiceover artists, audiobook producers and developers who need the best available speech engine should still start here, and should budget above the tier that looks sufficient. Anyone who mainly wants music or video, or who needs the same line to come back identical every time, is better served elsewhere.

Community

Discussion

Join the discussion and share your perspective.