Buying a synthetic voice looks like a simple purchase until the first invoice and the first legal question arrive. One team wants forty minutes of narration for a compliance course and needs a license that lets the files leave the building. Another is putting speech into a support product and counts characters per month. A third wants the founder's voice cloned for a podcast intro, and has not asked who else could ask for it. Those are three different products that share a landing page. This ranking sorts 12 of them for a company of 20 to 500 people by what can be checked before paying: which meter the bill runs on, which tier actually permits commercial use, whether recordings and scripts may train the vendor's models, and what stops a stranger from cloning a voice that is not theirs.
Visibility in this ranking can be paid for. Payment moves a vendor's position within the
shortlist; it never adds a vendor, and it never changes a word of the review. The largest vendors in AI voice generators
cannot hold places 1 to 3. How it works: placement disclosure ·
editorial process.
How we ranked these
Setup: choosing a voice, fixing pronunciations, proving consent
Getting a first file out takes minutes in a studio product and an afternoon of API keys in a developer one. The slow part is elsewhere. Product names, acronyms and people's surnames come out wrong on the first pass, so the useful question is how a team corrects them: a pronunciation dictionary, phonetic markup, or re-typing the script until it sounds right. Cloning adds a second task. Resemble AI may ask the person being cloned to record a spoken consent statement before a voice model is built, and Descript works from a consent statement too, while others only ask the customer to confirm permission with a checkbox. We read what each vendor documents on its own pages, and noted where custom voices need an application, as on Azure, or an enterprise contract.
The real price: characters, credits or downloaded minutes
Three meters are in use and none converts cleanly into another. Cloud APIs charge per million characters, which is the one honest unit: Amazon Polly lists $4 for standard voices and $16 for neural, Google lists $4 to $160 depending on voice type, and Deepgram bills $30 per million characters for Aura-2. Studio products sell credits or downloaded minutes, so a tier's headline allowance means little until it is converted into finished audio, and WellSaid's 20 minutes a month on Starter is a hard ceiling rather than a soft one. Free tiers are mostly test benches. ElevenLabs' free plan and Typecast's free plan both withhold or condition commercial use, so a company should price the first tier that licenses the audio, not the first tier that exists.
Getting your data out: audio files, voice models and the training clause
The audio itself is portable: a WAV or MP3 plays anywhere, and the only blockers are tiers that cap or restrict downloads. What does not travel is the voice. A cloned or custom voice lives inside the vendor's service as a model the customer cannot export, so leaving means recording again. The inbound direction needs more attention than the outbound one. Scripts are often confidential, and the voice recordings used for cloning are biometric data in several jurisdictions. Two vendors here state that customer material is not used to train their models, WellSaid and Resemble AI; five reserve the right, some with an opt-out form or a request flag; five say nothing definite on the pages we read. We recorded each position instead of assuming one.
Independence from the vendor: who holds the voice, and where it runs
Dependence takes two forms. The first is the voice: a branded or cloned voice has no second home, so a price rise or a shutdown means rebuilding it with another provider. The second is the infrastructure. Eleven of the 12 vendors here are headquartered in the United States by their published address, so for a European buyer the question is where processing happens and which legal entity is the controller. ElevenLabs, for example, names a Polish company as controller of voice data in its privacy policy. Resemble AI offers on-premises deployment on its Enterprise tier and Cartesia describes cloud, on-premise and on-device options, which are the only routes on this list that keep audio inside a company's own environment.
Who it's for: narration, product speech or cloned voices
Most disappointments come from buying the wrong kind. A training department that needs narrated modules does not want a latency-tuned streaming API; it wants a studio with a script editor and a plain license. A product team building an assistant has the opposite need and wants a published per-character rate and concurrency limits. A marketing group chasing a brand voice needs cloning and consent handling above all. Each entry below names the job it is built for. The ranking favors products a mid-sized company can buy from a price list, and, by rule, places the four platform-scale vendors, ElevenLabs, Google, Amazon and Microsoft, below the third position.
The 12 tools, reviewed
#1 WellSaid
Studio narration for training and marketing audio · United States · wellsaid.io
Commercial rights SOC 2 Closed model
WellSaid does one job, narration for training, product and marketing audio, and the paperwork around it is the cleanest in this group. The pricing page lists what each plan downloads per month and grants commercial rights from Starter, so the cost of a finished module can be worked out before signing up. The company says its closed-model architecture keeps scripts and voice samples from leaving the customer's environment or training external models, and states SOC 2 and GDPR compliance. That combination is what a compliance or learning team usually asks first, and it is why it sits at the top.
Where it falls short
Downloads are the meter and the caps are low: 20 minutes a month on Starter, 180 on Pro. API access and custom voices appear only through Enterprise, which is quoted, so a team that wants to wire narration into its own software cannot price that from the site. The free trial allows 3 downloaded minutes and no commercial use, which is enough to listen and not much more.
Wrong for
A product team that needs speech generated live inside an application, or a creator who wants a cloned voice on a $20 plan. Cartesia or ElevenLabs fit the first need; Typecast fits the second.
Pricing: Trial is free with 3 downloaded minutes a month; Starter is $19 a month or $120 a year; Pro is $49 a month or $396 a year; Business is $160 per user per month billed annually; Enterprise is quoted. (free trial)
Visit WellSaid →
#2 Typecast
Voice studio with emotion controls and avatars · South Korea · typecast.ai
Emotion controls Voice cloning Video editor
Typecast is the cheapest serious way into this category on the list. Generating and previewing audio costs nothing on any plan; credits are only spent when a file is downloaded, which lets a producer iterate on tone without watching a meter. Instant cloning is on the $5 Basic plan, emotion and intonation controls arrive on Pro at $29, and the same account includes a video editor and talking avatars. The pricing page is explicit that attribution is required on free-plan downloads and optional from Basic, so the license boundary is easy to find. For a five-person marketing team that is a good deal.
Where it falls short
The credit allowances are modest: Basic is described as about 35 minutes of download a month. Professional cloning needs Plus, and the API and enterprise terms are sold by inquiry rather than listed. We could not locate a readable terms or privacy page this session, so its training and retention positions are unverified.
Wrong for
A regulated company that needs a documented data-handling position, SSO or a deployment option before it buys. WellSaid or Resemble AI publish more on those points.
Pricing: Free is $0 with about 5 minutes of lifetime download credit; Basic is $5 a month, Plus $19, Pro $29 and Business $69, with discounted yearly billing; Enterprise and API are by inquiry. (free plan)
Visit Typecast →
#3 Resemble AI
Cloned voices with detection and watermarking · United States · resemble.ai
On-premises Consent check Deepfake detection
Resemble AI earns its place by treating misuse as a product problem. The privacy policy says customer recordings are used to deliver the contracted voice model and for nothing else, and that customer voice data does not train general-purpose models. A person being cloned may be asked to record a spoken consent statement, the platform sells detection and watermarking next to synthesis, and the Enterprise tier can be deployed on-premises. The homepage cites SOC 2 Type II, HIPAA and GDPR. For a bank, a broadcaster or a call center worried about voice fraud, that set of controls matters more than another stock voice.
Where it falls short
The pricing page is built around detection and platform seats, and it does not show a per-minute or per-character price for generated speech, so a team buying narration cannot cost it from there. Team starts at $350 a month with five seats. The library of ready-made voices is not what the product leads with, and it is a poor fit for a one-off voiceover.
Wrong for
A small team that needs a few narrated videos this month on a modest budget. Typecast or WellSaid are priced and packaged for that, and need no platform contract.
Pricing: Flex is $0 on pay-as-you-go; Team is $350 a month or $280 billed yearly; Business is $1,000 a month or $800 billed yearly; Enterprise is quoted. The page we read lists detection rates, not synthesis rates. (free plan)
Visit Resemble AI →
#4 ElevenLabs
Voice platform for narration, cloning and agents · United States · elevenlabs.io
Voice cloning API Large library
ElevenLabs is the product most buyers will have heard of, and its price ladder is the most granular on the list: ten thousand free credits, then five paid steps up to $990 a month for 6 million credits. Cloning starts on the $6 Starter plan and professional cloning on Creator, API access is on every tier and the commercial license starts at Starter. If voice quality and voice choice are the deciding factors, it belongs on the shortlist for a trial. It sits at fourth because the ranking puts dominant platforms below the third place, not because the product is weak.
Where it falls short
Credits are not minutes and the conversion changes with the model and feature in use, so a monthly bill is hard to predict. The free plan carries no commercial license. The privacy policy says audio, text and video may be processed to train its models, with practices meant to dissociate them from identity, which is a position some legal teams will need to review. Cloning slots are capped per tier.
Wrong for
A company whose policy bars vendor training on uploaded material, or one that wants a fixed price per finished minute. WellSaid states its closed-model approach and bills in downloadable minutes.
Pricing: Free has 10,000 credits a month with no commercial license; Starter is $6, Creator $22, Pro $99, Scale $299 and Business $990 a month, each cheaper billed yearly; Enterprise is quoted. (free plan)
Visit ElevenLabs →
#5 Speechify
Reader app, voiceover studio and speech API · United States · speechify.com
Studio Developer API Dubbing
Speechify matters here for its pricing transparency. The Studio page shows a free plan, then $100 or $300 a year with credit allowances and a statement of when commercial rights begin, and the API page lists per-character rates down to $6 per million, with 500,000 free characters a month. A business that needs voiceovers and also a speech API can get both from one vendor on published terms. The studio allowances are large for the money, and the annual pricing is easy to budget.
Where it falls short
The name covers three products with three price pages, and the consumer app's terms say most services other than Voice Over Studio are not intended for commercial use, which a buyer can miss. Studio's free plan carries no commercial rights. Annual billing for Studio makes a single month of use a poor value. We found no clear training statement for the voice products.
Wrong for
A company that wants one contract and one meter, or a buyer who needs to host the voice in its own environment. A single cloud API from Google, Amazon or Microsoft covers that better.
Pricing: Studio has a free plan with 600 credits, Starter at $100 a year and Creator at $300 a year; the API starts free with 500K characters, then $10, $99 and $499 a month; the reader app Premium is $29 a month. (free plan)
Visit Speechify →
#6 Cartesia
Streaming speech models for voice products · United States · cartesia.ai
Streaming TTS On-device Voice agents
Cartesia publishes the details a developer needs to cost a voice feature: credits, an approximate minute count for each tier, and concurrent request limits that rise from 2 on Free to 15 on Scale. The commercial license and instant cloning begin at the $5 Pro plan, seats are unlimited on every plan, and the vendor says the same models can run in the cloud, on-premise or on-device. For a team shipping speech into software, that candor about limits is useful and rare. It lands in sixth place because it is a component, not a narration tool, and the training terms need a decision.
Where it falls short
The terms say inputs, outputs and interactions may be used to train its models, with an opt-out request form that applies only going forward. There is no narration editor, so a non-technical team has nothing to work with. Professional voice cloning needs the $49 Startup plan, and localizing a voice to another accent costs extra credits.
Wrong for
A communications team that wants to paste a script, pick a voice and export a file. Use WellSaid or Typecast, which give an editor, and keep Cartesia for product features.
Pricing: Free is $0 with 20K credits a month; Pro is $5 with 100K credits; Startup is $49 with 1.25M; Scale is $299 with 8M; Enterprise is quoted. The page converts plans into about 27, 133, 1,667 and 10,667 minutes of speech. (free plan)
Visit Cartesia →
#7 Descript
Text-based audio and video editor with AI voices · United States · descript.com
Text-based editing Overdub Podcast editing
Descript is the answer when the real job is editing a recording and synthesized speech is only a patch for it. Fixing a misspoken word by typing it, generating a replacement in the speaker's own cloned voice and exporting the result in one tool is a workflow no pure generator gives. The paid tiers include text-to-speech with custom voice clones and stock speakers, and the Business tier adds dubbing in 30 or more languages. Its voice cloning works from a consent statement read by the speaker, which gives the practice a documented basis.
Where it falls short
Its terms allow Descript to use inputs and outputs to train and improve its models unless the customer opts out, and say third-party providers are contractually barred from doing the same. Pricing runs on media hours and AI credits, so a user who only needs narration pays for an editor.
Wrong for
A team that needs hundreds of narrated lessons from scripts, with no recording to edit. WellSaid is built for that work, and a cloud API is cheaper at that volume.
Pricing: Free plan with limited AI tools; Hobbyist, Creator and Business tiers are priced per month with media hours and AI credits; Enterprise is quoted. The page showed monthly and annual figures inconsistently, so no number is given. (free plan)
Visit Descript →
#8 Deepgram
Speech APIs for voice agent builders · United States · deepgram.com
Aura TTS Usage-based Speech to text
Deepgram belongs here for the buyer who already transcribes with it or plans to build a voice agent that listens and speaks through one vendor. Its text-to-speech rates are published to four decimals, there is no minimum or expiry on the pay-as-you-go credit, and the free $200 credit funds a real pilot rather than a demo. The terms add a per-request parameter to opt out of model training, which is more control than most APIs give. Concurrency limits are listed on the pricing page.
Where it falls short
Speech recognition is the core business, and text-to-speech is a younger line with a small voice catalog compared with studio products. There is no editor, no cloning on the pricing page we read and no ready-made narration workflow. The default position in the terms is that customer content improves its models, and the opt-out has to be sent with each request.
Wrong for
A content or learning team that needs finished voiceovers, not an API. Typecast and WellSaid supply the editor, voices and license that Deepgram leaves to the customer to assemble.
Pricing: Pay-as-you-go includes $200 free credit, then Aura-2 is $0.030 per 1,000 characters and Aura-1 $0.015; the Growth plan starts at $4K a year and lowers these rates by about 10 percent. (free trial)
Visit Deepgram →
#9 Google Cloud Text-to-Speech
Per-character speech API on Google Cloud · United States · cloud.google.com
Cloud API Free quota Custom voice
For a company that already runs workloads on Google Cloud, this is the least friction route to speech in an application. The pricing page lists every voice type with its rate and its monthly free quota, from $4 per million characters for Standard and WaveNet up to $160 for Studio, so cost is a spreadsheet exercise. Instant custom voice is a published add-on at $60 per million characters. Billing sits on the existing cloud invoice, and no separate vendor review is needed if Google is already approved. As a platform vendor it is held below the top three by rule, and sits ninth.
Where it falls short
It is an API, not a studio: nothing here lets a non-engineer script, preview and export narration. Voice type naming has accumulated across generations, so choosing among Standard, WaveNet, Neural2, Studio and Chirp takes some testing. Characters are billed even for text that is regenerated, and a published statement on the use of customer content for training was not in the pages we read.
Wrong for
A marketing or training team with no developers. WellSaid or Typecast give them an editor and a license in one subscription, without a cloud account.
Pricing: Pay-as-you-go per character: Standard and WaveNet US$4 per million, Neural2 US$16, Chirp 3 HD US$30, Studio US$160, instant custom voice US$60; monthly free quotas of 1 to 4 million characters by voice type. (free plan)
Visit Google Cloud Text-to-Speech →
#10 Amazon Polly
Pay-as-you-go text to speech on AWS · United States · aws.amazon.com
AWS Pay-as-you-go Speech marks
Polly is the cheapest way on this list to turn a large volume of text into speech, and the arithmetic is simple: characters in, a rate per million, nothing else. Standard voices cost $4 per million characters with a free 5 million a month, and replaying cached speech is not billed again. Speech marks metadata helps with synchronizing text highlighting and lip movement. Its FAQ says the customer keeps ownership of processed content. Platform vendors rank below the third place by rule, and this one sits tenth.
Where it falls short
The cheapest rate applies to standard voices only, and the generative and long-form voices cost $30 and $100 per million characters. The FAQ says Polly may store and use text inputs to improve AWS services, with an organization-level opt-out policy. There is no voice editor and no consumer-style cloning in the pages we read, so output tuning goes through SSML.
Wrong for
A team that needs expressive narration or a branded cloned voice with no engineering help. ElevenLabs, Typecast or Resemble AI are closer to that need.
Pricing: Per million characters: Standard $4, Neural $16, Generative $30, Long-form $100. Free monthly allowance is 5 million standard characters; neural, long-form and generative quotas last 12 months. (free plan)
Visit Amazon Polly →
#11 Azure AI Speech
Neural voices and custom voice on Azure · United States · azure.microsoft.com
Azure Custom voice HD voices
Azure AI Speech suits organizations whose identity, billing and security reviews already run through Microsoft. It gives neural and HD voices, SSML control, a monthly free tier of half a million characters and commitment tiers for predictable volume. Its custom voice is offered under limited access for responsible use: the customer applies, and only after approval can it build a professional voice. That gate is slower than a checkbox, and it is also the most deliberate consent control of the three platforms. As a platform vendor it takes the eleventh place.
Where it falls short
Pricing varies by region and currency, and the page is long enough that a first reading does not produce one number. Custom voice cannot be self-served without approval. The speech portfolio lives inside Foundry tooling alongside many other services, and there is no simple studio for a non-technical narrator.
Wrong for
A small team without an Azure subscription, or one that wants a flat monthly fee and a single login. Typecast or WellSaid cost less in time and in administration.
Pricing: Free tier F0 includes 0.5 million neural characters a month; paid use is per million characters by voice type, with commitment tiers for volume. Rates vary by region and currency, so none are quoted here. (free plan)
Visit Azure AI Speech →
#12 Fliki
Text to voice and video for content creators · United States · fliki.ai
Script to video Voice cloning Creators
Fliki is the lightest tool here: paste an article or script, choose a voice and receive a narrated video with stock footage. The tiers are clear about what each unlocks, with one voice clone and commercial rights on Standard, three clones and the API on Premium, and a watermarked free plan that limits videos to five minutes at 720p. For a content marketer converting blog posts into short videos, it spares the cost of a separate video editor. It sits last because its pricing and policy pages gave us the least to verify.
Where it falls short
Credits blur the real cost, and the plan prices did not render for us. The privacy policy names outside AI processors, Runware, fal.ai, D-ID and OpenRouter, that receive face media, so the stack is rented from several parties. It states a no-training position for face media but we found none for voices. It is a creator tool, not a compliance-grade product.
Wrong for
A company that needs SSO, an audited security posture or consistent narration across hundreds of training modules. WellSaid fits that; Fliki fits individual creators.
Pricing: Free has 3 credits a month with a watermark; Standard has 180 credits and commercial rights; Premium has 600 credits and API access; Enterprise is quoted. Tier prices did not render on the page we read. (free plan)
Visit Fliki →
What the data says about this market
Counting from this list, the category is more open on price than on policy. Ten of the 12 vendors publish figures we could read, the exceptions being Fliki, whose tier prices did not render, and Descript, whose monthly and annual prices came back inconsistent. Ten offer a free plan of some kind, two a free trial. On the policy side only WellSaid and Resemble AI put a no-training statement in writing, five vendors (ElevenLabs, Cartesia, Descript, Deepgram, Amazon Polly) reserve some right to use customer material, and the remaining five left the question unanswered on the pages we read. A buyer comparing entry prices alone would miss the larger differences.
Eleven of the 12 are headquartered in the United States by published address; the twelfth, Typecast, is in Seoul. The list splits cleanly by customer. Four vendors, Amazon Polly, Google Cloud, Azure AI Speech and Deepgram, sell speech by the character or by usage, and three of them publish rates in dollars per million characters. Another group sells studio subscriptions metered in credits or downloaded minutes, and a third, Cartesia and ElevenLabs among them, straddles both and adds voice agents. Several pages now talk about voice agents and speech-to-text next to text-to-speech, which tells a buyer that the product being sold is increasingly a speech stack rather than a narration tool.
Voice cloning is the feature that has changed the risk profile. Instant cloning appears on entry tiers at Typecast (from $5 a month) and Cartesia (from $5 a month), and ElevenLabs lists it from $6, which means a cloned voice is now a commodity feature rather than an enterprise one. The compliance controls have not become commodities in the same way. Resemble AI sells deepfake detection and watermarking next to synthesis, which is a sign of where the sector is heading. The World Bank series this site tracks shows ICT services rising from 9.1 percent of world service exports in 2013 to 14.48 percent in 2023 (EU: 10.58 to 17.55), and synthetic speech bought by card from a vendor in another jurisdiction is a small example of it.
The 12 ranked vendors, counted
- Headquarters by region: North America 11, Asia-Pacific 1
- By country: United States 11, South Korea 1
- Pricing model: Credits / month 4, Per million characters 3, Downloaded minutes / month 1, Media hours + AI credits 1, Per 1,000 characters 1, Per product tiers 1, Platform tiers + usage 1
- Free option: Free plan 10, Free trial 2
Counted from the 12 vendors on this page. More in our market data.
For the wider market behind AI voice generators, read our report The Global Shift to ICT Services,
or browse all industry reports.
Questions and answers
What is the best AI voice generator in 2026?
For most companies it is WellSaid, because it sells narration with a plain commercial license, publishes its tiers, says in writing that customer scripts do not train external models and states SOC 2 and GDPR compliance. Typecast, second, is the better pick for a small team that wants cloning, emotion controls and a video editor from $5 a month. Resemble AI, third, suits a business that needs on-premises deployment, consent verification and deepfake detection. If the need is speech inside a product rather than narration, start at Cartesia or one of the cloud APIs lower down.
Can I use audio from a free AI voice generator plan commercially?
Usually not without checking. ElevenLabs' terms limit free-plan users to non-commercial purposes and its pricing page lists a commercial license from Starter upward. Typecast requires attribution on everything downloaded on the free plan and offers an optional one from Basic. WellSaid's trial lists no commercial rights, and Speechify Studio's free plan states there are no commercial usage rights. Cartesia's commercial license starts at its $5 Pro plan. Cloud APIs from Google, Amazon and Microsoft differ: they meter usage rather than license tiers, so free quotas are about volume, not permission.
Do AI voice generators train their models on my scripts and recordings?
It depends on the vendor. WellSaid says its closed-model design keeps scripts and voice samples from training external models, and Resemble AI states it does not use customer voice data to train general-purpose models. ElevenLabs' privacy policy says it may process audio and text to improve its models using practices designed to disassociate content from identity. Cartesia and Deepgram offer an opt-out, Descript's terms allow training unless a customer opts out, and Amazon Polly points to an AWS Organizations opt-out policy. For the rest, treat silence as a question to put in writing.
How much does an AI voice generator cost per hour of audio?
There is no single figure, since the meters differ. An hour of speech is roughly 55,000 to 60,000 characters of text, so at Amazon Polly's $16 per million characters for neural voices the raw cost is under a dollar, and Deepgram's Aura-2 at $30 per million is under two dollars. Studio products cost more per hour but include the editor and a license: WellSaid's Pro plan includes 180 downloaded minutes a month. Check what a regeneration costs, because first drafts are rarely final and credit-based plans count every attempt.
Is it legal to clone a voice with these tools?
Only with the speaker's permission, and the vendors differ in how they enforce it. Resemble AI may require a recorded consent statement before a model is built, and Descript's terms require the consent of anyone whose voice is used to train an AI voice. ElevenLabs' privacy policy notes that voice data can count as biometric data under applicable law. The tool's checkbox does not transfer the legal duty: the customer still needs a written release, and in the EU a data protection assessment, before cloning an employee, contractor or voice actor.
What is the difference between a text-to-speech API and a studio product?
A studio product is an application: a script editor, a voice picker, a timeline, and an export button, sold per seat or by credits. An API is a service your own software calls, billed per character and judged on latency and concurrency. Amazon Polly, Google, Azure, Deepgram and Cartesia are API-first. WellSaid, Typecast, Descript and Fliki are studio-first. ElevenLabs and Speechify sell both, with separate pricing, which means two contracts if a company wants both.
Which AI voice generators can run on our own infrastructure?
Few can. Resemble AI lists on-premises deployment on its Enterprise tier and describes air-gapped options on its homepage. Cartesia describes the same models running in the cloud, on-premise and on-device, with inference kept in-region. The rest are multi-tenant cloud services. For the three hyperscalers the practical control is the region you pick, not the location of the model. If audio cannot leave your network for regulatory reasons, ask the vendor for that deployment in writing and see whether it is a product or a project.
Do AI voices work in languages other than English?
Yes, but coverage varies sharply and is easy to overstate. Speechify's reader app lists 60+ languages, Fliki lists 80+ languages across its voices, and Typecast's site cites 35+ languages. Google, Amazon and Microsoft cover many locales but quality differs by voice type. Cartesia notes that localizing a voice to another accent costs extra credits. Test the languages you need with your own script, including names and numbers; a vendor's language count says little about how natural one particular language sounds.
Can I put an AI voice on a phone line or an in-app assistant?
Yes, and for that job you want an API with published concurrency limits, not a studio. Cartesia lists concurrent requests by plan, from 2 on the free tier to 15 on Scale, and Deepgram lists concurrency of up to 45 for text-to-speech on pay-as-you-go. Speechify's developer API and ElevenLabs also sell streaming access. Model the peak number of simultaneous calls rather than the monthly volume, since a concurrency cap shows up as a failed call, whereas a character quota only shows up on the bill.
Will the voice sound the same if we switch vendors?
No. A stock voice is the vendor's property and a cloned voice is a model that stays on the vendor's platform, so switching means recording samples again and accepting a slightly different result. Keep the original studio-quality recordings, the written consent and the script archive so the rebuild takes days, not weeks. Audio files you have already generated keep working, subject to the license of the tier that produced them. Avoid building a brand around a stock voice that other customers can pick too.
Is Play.ht still an option?
We could not confirm it. The play.ht domain did not resolve when we checked it for this ranking, so we left it out rather than list a product whose status we could not verify. LOVO and Murf are also not covered here, because we could not read their current pricing pages. All three are products a buyer will hear about, so verify their current terms with the vendor, and apply the same checks as for the entries here: the license tier, the training clause and the consent step.