HappyHorse

Alibaba's open-source AI video generator

In-depth Report

  • HappyHorse 1.0 is an open source AI video generation model launched by Alibaba ATH (Alibaba Token Hub) Innovation Division. In early April 2026, it appeared on the Video Arena rankings of Artificial Analysis, the world's authoritative AI evaluation platform. When it debuted, it won the first place in the two core evaluations of text generation video and image generation video, with an ELO score as high as 1357-1406 points, becoming a landmark product in the field of open source video generation. This model was developed by a team led by former Kuaishou Vice President Zhang Di. Its core advantage lies in its native audio and video synchronization generation capabilities under a unified architecture of 15 billion parameters. It supports accurate lip synchronization in 7 languages. It is still in the internal testing stage and has not yet been fully opened to public testing.

  • HappyHorse 1.0 was developed by the Future Life Lab of Alibaba ATH Innovation Division, which is affiliated with Alibaba Taotian Group. The project leader is Zhang Di, who previously served as the vice president of Kuaishou and the technical director of Kling AI. He has profound technical accumulation and industry knowledge in the field of video generation. On March 16, 2026, Alibaba officially established the Alibaba Token Hub (ATH) business group, which is directly responsible for CEO Wu Yongming and positioned as a next-generation AI content creation platform. HappyHorse is one of the first exploration results of this business group. The ATH AI Innovation Division has launched an exploration plan for interactive methods in the AI ​​era, and more products will be launched in the future. At the level of technical cooperation, HappyHorse is supported by Sand.ai (Autoregressive World Model) and the GAIR Laboratory of Shanghai Intelligent Computing Research Institute. This model evolved based on the daVinci-MagiHuman project that was open sourced in March 2026, but underwent large-scale upgrades in architecture and performance, ultimately forming today's 15 billion parameter unified multi-modal large model.

  • The release of HappyHorse coincides with a period of fierce competition in the AI ​​video generation track. Since 2026, major products such as OpenAI's Sora 2, ByteDance's Seedance 2.0, Kuaishou's Kling 3.0, and Google's Veo 3 have been launched one after another. Video generation is becoming one of the most watched directions in the AI ​​field. At this point in time, HappyHorse directly climbed to the top of the rankings in an airborne manner. There was no press conference or pre-marketing. It purely relied on its technical strength to trigger a shock in the industry, which reflected Alibaba's deep accumulation in the field of AI video and its strategic intention to suddenly exert force.

  • HappyHorse 1.0 adopts a unified single-stream Transformer architecture with 15 billion parameters and a total of 40 self-attention layers. The core innovation of this architecture lies in the unified multi-modal design: text tokens, image latent variables, video frames, and audio waveforms are packaged into a sequence for joint denoising, without the need for cross-attention modules and no external audio models. The middle 32 layers share parameters among all modalities, and the first and last 4 layers use modality-specific projections to achieve truly integrated audio and video generation. This model applies DMD-2 distillation technology to reduce the denoising steps from the traditional dozens of steps to only 8 steps, greatly improving the inference efficiency. On an NVIDIA H100 single card, the 256p preview mode can generate a 5-second video in about 2 seconds, and the full version of 1080p with synchronized audio can generate a 5-second video in about 38 seconds.

  • In terms of generation capabilities, HappyHorse supports text-generated videos, image-generated videos, multi-camera narratives and cross-scene character consistency. Users can generate complex scenes through simple text descriptions, including camera movement, lighting conditions, character actions, etc. For static images, the model can convert them into dynamic videos, maintaining subject consistency. Synchronous generation of native audio and video is the core highlight of this model. Video and audio are synchronized in a single generation process, rather than dubbing afterward. The model supports accurate lip synchronization in 7 languages, including Mandarin, Cantonese, English, Japanese, Korean, German and French, with extremely low WER word error rates. This solves the long-standing pain points in the field of AI video of silent videos or out-of-sync audio and video. In terms of output specifications, the model supports multiple aspect ratios of 16:9, 9:16, 4:3, 3:4, 21:9, and 1:1. Each time it generates 5-12 seconds of video, it supports up to 1080p resolution and can be upgraded to 2K. For commercial use, each video enjoys 100% commercial copyright.

  • According to the official website, more than 10,000 creators around the world are using HappyHorse 1.0. Judging from user feedback, the multi-camera narrative function is considered disruptive, and the native audio generation shocks users. The 2K movie-level image quality has reached professional production standards, and the generation speed is about 30% faster than similar tools. New users can get points for free to experience all functions without binding a credit card.

  • HappyHorse offers three levels of subscription packages. The basic version costs US$11.90 per month, provides 540 points, and can generate about 54 videos; the professional version costs US$39.90 per month, provides 2040 points, and is suitable for frequent creators; the studio version costs US$99.99 per month, provides 6000 points, and supports high-intensity commercial use. Points are reset monthly and cannot be accumulated. As an open source product, HappyHorse provides a complete open source version, including the base model, distillation model, super-resolution module and inference code. Users can fully self-host and run locally, which is attractive to commercial users with privacy requirements or the need for deep customization.

  • Judging from public information, users’ feedback on HappyHorse is generally positive. Positive comments are mainly concentrated in several aspects: the simultaneous generation of native audio and video is considered an industry breakthrough, solving the past limitations of AI videos that can only generate pictures and post-dubbing; the multi-lens narrative function supports the generation of complex story scenes, which is recognized by creators; the generation speed and image quality are better than most competing products; the completely open source and commercial model reduces the cost of use for creators. The technology community pays great attention to HappyHorse's architectural innovation. The unified single-stream Transformer design with 15 billion parameters achieves true audio and video integration while ensuring performance, and is considered an important technological breakthrough in the field of video generation.

  • In the latest evaluation of Artificial Analysis Video Arena, HappyHorse 1.0 achieved impressive results: text-generated video without audio ranked first with an Elo score of 1333-1357, about 60 points ahead of the second-placed Seedance 2.0; image-generated video set a record high in the platform's history with 1391-1406 points; it also stably maintained its second place in the audio-containing mode. This achievement makes it the most powerful open source AI video generation model in the world. Compared with closed source competitors, HappyHorse has unique advantages in open source ecology and self-hosting capabilities. Compared with closed source models such as Seedance 2.0, Kling 3.0, and Veo 3, its core difference lies in its architecture design that is completely open source, can be deployed locally, and has native audio and video synchronization.

  • Media reports generally use words such as dark horse, top-selling, and subversive to describe the release of HappyHorse. Chinese technology media such as Caijing, 36kr, Aifaner, and Geek Park have all reported on this. Industry insiders commented that the competition in video generation has just begun, and HappyHorse has just opened the first page of a new chapter, implying that the competition in the AI ​​​​video track is far from over. On April 30, 2026, HappyHorse will officially open its API and will enter a new stage of commercial operation. The outside world is generally concerned about its API pricing and openness, which will become an important window for judging its commercialization strategy.

  • Currently, HappyHorse is still in the internal testing stage and has not yet been launched on a large scale. According to IT House’s report, many official website addresses circulating on the Internet are not officially certified versions, and we still need to wait for the official product to meet everyone. Early users may face problems such as unstable functionality and long queue times.

  • Although an open source version is provided, the balance between open source and commercialization is a challenge. If the open source version is too functional, it may affect paid subscription revenue; if the performance gap between the open source version is too large, it may affect the participation of the developer community. How to find a balance between open source ecology and commercial monetization is a long-term issue faced by HappyHorse.

  • AI video generation technology has the potential risk of being abused for deepfake. As the quality of generation improves, how to prevent technology from being used in scenarios such as fake celebrities and fake news requires joint efforts by platforms, regulatory agencies and users.

  • HappyHorse is particularly suitable for the following users: professional video creators and short video practitioners, content creators who need to quickly generate product promotions, social media short films, concept previews and short narrative videos, enterprise users who give priority to open source self-hosting capabilities, projects with high requirements for audio and video synchronization quality and multi-language lip synchronization, as well as developers and researchers who study and learn video generation technology.

  • For simple scenarios with extremely limited budgets and low quality requirements, the paid version may not be as cost-effective as free tools; for projects that require long videos longer than 12 seconds, the current version may not meet the needs; for commercial advertisements that require specific brands or characters to appear, you need to pay attention to compliance requirements.

  • In the closed-source solution, Seedance 2.0 ByteDance, Kling 3.0 Kuaishou, and Veo 3 Google are the main competing products. Among the open source solutions, Wan 2.6 and others can be used as a reference choice, but there is a gap between the overall performance and HappyHorse.

  • HappyHorse 1.0 is Alibaba's blockbuster product in the field of AI video generation. Its native audio and video synchronization generation capabilities under a unified architecture of 15 billion parameters have a clear leading edge at the technical level. It topped the list in authoritative global reviews on its debut, proving its technical strength. As a completely open source product, it provides new options for content creators, especially attractive to users who require local deployment and deep customization. The product is currently still in the internal testing stage, and the API will be officially opened on April 30, 2026. At that time, its commercialization strategy and API pricing will become the focus of the industry. Competition in the AI ​​video generation track is fierce. Whether HappyHorse can maintain its technological leadership after public testing and commercialization requires further observation.

User Reviews

  • 头像
    happyfrog634
    就盼着正式开放 API 了,现在还在灰度,每次想批量跑还得排队等名额。

  • 头像
    Lauren190
    复杂流体和吊臂长跟拍还是 Kling 3.0 更稳,HappyHorse 在对话驱动、多语言唇同步、电商面料纹理这些场景是真的强。九张参考图输入角色控制很稳。一句话,别拿它干不擅长的事,找准场景性价比拉满。

  • 头像
    AnnHoward_X
    唇音同步准,但四句快对话塞进 15 秒有点挤,音乐场景里乐器声偶尔和手不同步,希望后续修。

  • 头像
    Gabriel_Anderson04
    1.1 版皮肤纹理真实多了,毛孔细纹都出来了,不再是那种数字光泽感。

  • 头像
    DebbieWard
    对比了下价格,HappyHorse 官网折扣价 720P 才 0.44 元/秒、1080P 0.78,Seedance 2.0 要 1 元和 2.48。对中小团队来说这个差价积少成多很可观。企业客户走阿里云 API 还没门槛,字节之前要预缴千万级,这波阿里挖客户很猛。

  • 头像
    VMillerX
    看到 GitHub 上说能下权重赶紧跑,结果发现是假的,官方早就说了闭源没权重,提醒大家别上当。

  • 头像
    Samantha.Lewis_2024
    排队比 Seedance 舒服多了,没怎么等。

  • 头像
    Alexander.GonzalesX
    说实话它最适合当广告和短剧里的「中间镜头」——人物情绪、生活场景、B-roll 空镜、转场这些。过去靠外拍靠模特靠场地,现在一条 prompt 加几块钱解决。但想靠它替代导演还早,物理、音频、长时一致性都有边界。

  • 头像
    Jessica.Kim_Plus854
    做护肤品竖屏广告那条绝了,瓶身金色文字全程清晰没乱码,手模肤质高光都真实,以前得租影棚请手模,现在一句 prompt 搞定。

  • 头像
    sz1lhwiyf7
    音画同步是真香,背景蝉鸣和画外音都自己生成了,不用再后期配音。

  • 头像
    Diane_Hicks369
    烧了一万积分测了快一百条视频,最打动我的是光影和微表情。爷爷低头看屏幕时眯眼抿嘴那种老年人的笨拙感,还有雨夜出租车窗外霓虹扫过脸的安静张力,角色一致性稳得离谱。但物理偶尔翻车,穿模和人物突然消失还是会有。

  • 头像
    WHicks_2021514
    在千问 App 里试了下,输入巴黎街头那段,镜头跟着女主转男主的时候转场挺自然的,光影也有质感,整体超出预期。

  • 头像
    Anthony.Williams_Max
    灰测第一天就去官网薅了 66 积分,生成 5 秒才花 32 积分,价格比我想的友好。

  • 头像
    DRogers36995
    欢乐马这波属实惊艳了。

  • 头像
    Isabella_Torres520
    试了一下HappyHorse,确实牛。文字转视频比我之前用的Kling快太多了,关键是生成的视频里人物动作自然,不会出现那种诡异的肢体扭曲。

  • 头像
    Alan_WoodSr8
    唯一的问题就是积分不太够用,基础版一个月540积分也就生成50多条视频,性价比一般般。

  • 头像
    JHicksII
    阿里这次真的杀疯了,150亿参数开源随便用,香!

  • 头像
    SarahJenkinsIII
    用了三天,说下感受:音视频同步确实强,但是生成速度比官方说的38秒要久一点,可能是我网络问题?

  • 头像
    lazyduck280
    多镜头叙事功能是认真的吗?太好用了!描述一个场景就能自动切好几个镜头,省了我不少功夫。

  • 头像
    Jerry_MurrayK
    本地部署跑了一下,8卡H100出片很快,就是显存有点吃紧,建议32G以上显存。

  • 头像
    Jesse.Evans168
    比Seedance 2.0便宜太多了,而且还能自己部署,赞。

  • 头像
    Nicholas986
    还没用上的朋友赶紧去申请内测,真的香!

  • 头像
    George.Ramirez0
    真的绝了原生音频生成,之前用别的工具还得自己配音,现在一次搞定。

  • 头像
    黎月
    7种语言唇形同步也太卷了,以后做跨境内容方便多了。

  • 头像
    MarkHendersonSr511
    看了评测视频,ELO 1357不是吹的,画面流畅度和提示词还原度都很顶。

  • 头像
    CoinAnalystH_enderson
    画质可以的,1080p输出比很多付费工具清晰多了。

  • 头像
    MButler_2024_271
    开源免费不用白不用,已经在本地跑起来了。

  • 头像
    烟雨745
    唯一担心的是4月30号开放API后会不会涨价...

  • 头像
    JScott_2021
    yyds!屠榜了!

  • 头像
    whitebird444
    请求官方出移动端APP!现在只能在电脑上玩太不方便了。