ElevenLabs v4

ElevenLabs 预告的第四代 AI 语音模型,主打能低语、能吟唱、带情绪与口音的人类级表达

In-depth Report

  • ElevenLabs v4 is the fourth-generation speech synthesis model announced by this AI audio company with backgrounds in London and Warsaw. It is currently in a "public preview and has not yet been officially released." It was first demonstrated by founder Mati Staniszewski at ElevenSummit 2026 (Warsaw) and focuses on more natural human-level expression - the ability to whisper, chant, and speak any text with emotion and accent. What needs to be made clear is that as of July 2026, v4 has not yet been opened for public testing, and the flagship model that will actually run online is the Eleven v3 released in June 2025. Therefore, this report treats v4 as "the next generation direction that has been officially announced and previewed". The current status and pricing of the platform are based on the currently available version.

  • Founded in 2022 by Mati Staniszewski and Piotr Dąbkowski, ElevenLabs has its technical roots in Poland and is headquartered in London. It is one of the most representative generative audio unicorns in Europe. The company started from the judgment of "making AI voice no longer like a machine" and believed that the communication layer is the bottleneck of AI implementation, rather than pure "intelligence" itself. In a few years, it has grown from a single text-to-speech tool to a full-stack audio platform covering speech generation, dubbing, music, speech cloning, real-time transcription and speech intelligence. The customer list includes Disney, Epic Games (Darth Vader voice in Fortnite), Duolingo, as well as large-scale enterprise-level deployments such as Deutsche Telekom, Nvidia ACE, and the Greek Government Travel Assistant.

  • The current public information on v4 comes from the keynote demo in Warsaw. Staniszewski emphasized on the spot that "this model has not been released yet, this is an exclusive display." What the audience heard was a two-person conversation about Warsaw - the voice can switch accents, whisper, carry emotions, and even sing. According to the official statement, what v4 wants to solve is not "clearer reading", but "more like human communication": controllable emotions, clear intentions, natural accents and pauses. The same event also previewed a new dubbing system code-named D2. It no longer only generates speech based on text, but reads the emotions and performances of the original film, and transfers this "spirit" across languages, directly attacking the old problems of traditional AI dubbing "expression tablets". Compared with Eleven v3, which is already in use, v3 will be fully open on March 14, 2026. The core selling point is Audio Tags - write inline instructions such as [whispers], [excited], and [sighs] directly in the script, and the model will adjust the interpretation at the phrase level without repeated re-recording; it also supports multi-speaker Dialogue Mode, more than 70 languages, and the upper limit of a single request is about 3,000 characters. Further down, Multilingual v2 focuses on stability and consistency, Flash focuses on 75 milliseconds ultra-low latency, and Turbo focuses on high cost performance, forming a model matrix divided according to "quality/delay/cost". If v4 is launched as scheduled, it will most likely be above the expressive power of v3, pushing "naturalness" and "controllable emotions" to a higher level.

  • ElevenLabs pricing is open to the same pool of credits, with text-to-speech priced per character (approximately 1 character = 1 credit for a high-quality model, approximately 1000 characters for 1 minute of audio). The current public tiers are: Free version with 10,000 points per month; Starter with 6 USD/month, 30,000 points; Creator with 22 USD/month (11 USD in the first month), 121,000 points and unlocking professional voice cloning; Pro with 99 USD/month, 600,000 points, API can output 44.1kHz PCM; Scale with 299 USD/month, 1.8 million points, 3 work seats; Business with 990 USD/month, 600 10,000 points, 10 seats; enterprise version is negotiable. Other product lines such as music, sound effects, voice changing, and dubbing are deducted from the same point pool at different rates. For example, music is about 900 points/minute, and dubbing is about 2,000–10,000 points/minute. There is also a free 12-month, 33 million-character grant program for start-up teams. The specific billing for v4 has not yet been announced, third-party tracking sites speculate that it is about $0.30/thousand characters, but this is an unconfirmed rumor.

  • Sound of the existing v3 has been generally positive. Practitioners of long-form audiobooks, game dubbing, and film and television localization recognize its improved expressiveness. In particular, Audio Tags allow a single person to adjust the emotions of multiple characters, eliminating the need for traditional dubbing directors. In terms of blind test preferences, v3 has improved by 72% compared to the early v3 Alpha. Negative voices focus on two points: First, it is expensive, with a price of about US$100 per million characters, which is significantly more expensive than competing products in mass production scenarios; second, long texts need to be split into multiple requests, the automated workflow is somewhat cumbersome, and the quality of low-traffic languages ​​is still unstable. Regarding v4, users are still in the "expectation" stage. The community is more concerned about whether it can truly compress the mechanical feel and whether the price will continue to rise.

  • In the voice arena of Artificial Analysis (June 2026), the Elo of Eleven v3 is about 1178, firmly ranking in the first echelon of commercial TTS. However, it has been overtaken by Alibaba Fun-Realtime-TTS (about 1219), Google Gemini 3.1 Flash TTS (about 1214), Inworld, Cartesia Sonic 3.5, etc. in terms of raw naturalness, but the unit price is much higher than them. The industry consensus is that ElevenLabs wins in "ecological integrity + expressiveness + enterprise-level compliance", but "unit cost" is becoming its weakness. In terms of competing products, OpenAI TTS-1, MiniMax Speech, Cartesia, Inworld, and Suno (music-oriented) are all competing for the same piece of cake.

  • Sound cloning naturally treads deep waters. ElevenLabs has built-in content review, accountability, and watermark traceability mechanisms, and has repeatedly emphasized that "AI-generated audio should be identifiable." However, the risk of cloning technology being abused for deepfake, fraud, and impersonation is always there. At the commercial level, the copyright and authorization of training data are unavoidable topics - the company emphasizes that the music model is trained based on authorized data and can be used for commercial purposes, but the source of the voice side collection still triggers discussions from time to time. For ordinary users, the real risks are more simple: expensive points, re-run billing, opaque enterprise version terms, and the expected gap caused by major updates such as v4 that are "officially announced early and officially available later."

  • If you are a content creator who produces audio books, podcasts, game character dubbing, or film and television localization, currently using v3 directly with Audio Tags is the most pragmatic choice, and the expressiveness is sufficient. If you are a company building voice-enabled agents such as customer service, sales, and travel guides, you should look at ElevenAgents and Dubbing v2 instead of waiting for v4. For batch scenarios with sensitive budgets and large quantities and full management, it is recommended to conduct a horizontal blind test between ElevenLabs and MiniMax, Cartesia, and Alibaba TTS before making a decision. As for v4, it is recommended to pay attention to it as a "vane of the next generation direction" and wait for the official public beta, get the actual test samples and pricing before placing a bet, so as to avoid paying for undelivered capabilities in the preview stage.

  • ElevenLabs v4 is heading in the right direction - voice interaction should move from "speakable" to "human-like", but it is still just a preview sample on the stage in Warsaw. What really determines whether you want to use it today is the mature v3 ecosystem and the expensive points bill.

User Reviews

  • 头像
    Isabella.Howard_2020
    重跑也计费这点吐槽一下,生成错了重来还扣积分,批量出片时心在滴血。

  • 头像
    RebeccaLee
    希腊政府旅游助手那个 demo 挺实在的,能识图推荐行程还能看日历,比纯聊天机器人有用多了。

  • 头像
    JosephRamirez_7
    v4 若能控情绪和意图,游戏过场动画的旁白就能一次成型,不用再后期一段段调语气了。

  • 头像
    TIram
    竞技场里 v3 已经被阿里和 Gemini 反超了,但生态是真的全,从 TTS 到配音到智能体一把梭,小团队懒得接七八家 API。

  • 头像
    SatoshiFanBennett
    开发者资助计划真香,12 个月免费 3300 万字符,我们初创直接白嫖把智能体跑通了。

  • 头像
    BobbyMorrisX48
    v4 能吟唱这点我很期待,现在做游戏 NPC 哼个小调还得另接音乐模型,要是语音模型直接能带旋律就省事了。

  • 头像
    BLkra
    整体方向我看好,语音要从「能说」走向「像人」,ElevenLabs 这步棋下得对,就看定价别飘。

  • 头像
    陈萱
    Flash 的 75 毫秒延迟做实时对话基本够用了,v4 如果还能压一压,客服场景体验会再上一个台阶。

  • 头像
    RichardRuiz_Plus
    等不及 v4 了,现在先用 v3 的 Audio Tags 做有声书,单人就能调度几个角色的情绪,省了请配音导演的钱。

  • 头像
    MadisonPatelIII
    价格再不降,真的要被 MiniMax 和阿里系抢光了。我们公司一个月光配音就烧掉几百刀,老板已经在看替代方案。

  • 头像
    Mary_ButlerSr
    又是「预览」又是「独家展示」,正式公测怕是要等到年底,建议大家别为没交付的能力提前付费。

  • 头像
    DrMar'yanZaporozhec
    盲测里 v3 比早期 Alpha 提升 72%,说明他们打磨得狠,v4 正式版应该也不会让人失望。

  • 头像
    Patrick.EdwardsII
    Music 模型基于授权数据这点很重要,商用不踩版权雷,比某些来路不明的音乐生成靠谱。

  • 头像
    Bryan.RobertsSr
    在华沙那场演示里听到 v4 的样片时,鸡皮疙瘩起来了。那种带气声的低语和自然的停顿,确实是 v3 还差一口气的感觉。

  • 头像
    goldensnake135
    配音按分钟扣积分太肉疼了,千字配音动辄上万积分,大项目还是得咬牙上 Business 档。

  • 头像
    枫叶_13
    说真的,ElevenLabs 的克隆水印和溯源做得比大部分家都认真,这点我愿意多付点钱,至少不怕哪天被牵连进 deepfake 新闻。

  • 头像
    Andrea.Ramos
    关注 v4 很久了,建议把它当风向标看,真要投产还是先用成熟的 v3 生态,别赌预览阶段。

  • 头像
    Charles.Smith_2020
    D2 配音那个跨语言保留情绪的思路对,传统 AI 配音确实像念稿机器,希望 v4 时代能真正解决「表情平板」。

  • 头像
    RRuizZ170
    能不能先把长文本单次上限从 3000 字符提上去?自动化流水线天天切分请求,烦死了。

  • 头像
    Anthony_CoxJr
    英伟达 ACE 那次合作展示的多语言营销内容生成挺惊艳,声音自然度确实行业第一梯队。

  • 头像
    Madison.Vasquez007949
    用了两年 ElevenLabs 做播客,整体很稳,就是企业版条款不透明,想加席位和并发谈得头疼。

  • 头像
    Gary32
    低流量语种质量还是参差,泰语和越南语偶尔会冒奇怪的重音,正式项目我们还是切回英文再本地化。