Vidu

The first purely self-developed AI video generation model in China, supporting Wensheng videos, Tusheng videos and reference videos

In-depth Report

  • Vidu AI is the first purely self-developed AI video generation model in China, jointly released by Beijing Shengshu Technology Co., Ltd. and Tsinghua University. The platform focuses on converting text and images into high-quality dynamic videos while maintaining subject consistency. Vidu provides three core functions: Wensheng video, Tusheng video and Reference student video. The Reference student video function is the world's first video generation technology that supports multi-subject consistency. The platform claims to generate a video in 10 seconds and supports up to 1080p HD output. The Vidu Q3 version ranks first in China and second in the world in international authoritative evaluations, second only to xAI's Grok. It is suitable for social media content creators, professional film and television animation studios and commercial marketing teams. It is not suitable for users who pursue extremely low cost or completely free use.

  • Vidu AI was jointly developed by Beijing Shengshu Technology Co., Ltd. and Tsinghua University. It is the first purely self-developed large-scale long-duration, high-consistency, and high-dynamic video model in China. The operating entity is the Vidu brand, and the copyright belongs to Vidu®. Shengshu Technology was founded in 2024 and focuses on the research and development of multi-modal generative artificial intelligence technology. The founding team comes from the Artificial Intelligence Research Institute of Tsinghua University. In March 2025, Shengshu Technology announced that Vidu had achieved a new breakthrough in multi-modal technology and reached the world's top level in long video generation and cross-scenario connection. In April 2026, the Vidu Q3 version was released. Its core positioning is "born for dramas." Through the reference generation mechanism that can be referenced by everything, creators only need to provide a small number of reference images and natural language prompts to generate film and television-level special effects videos with one click. In the special evaluation list SuperCLUE-ComicShorts, Vidu Q3 ranked first with high scores and was evaluated as "the leading domestic product in the field of reference student videos."

  • Vidu AI provides three core video generation modes and rich auxiliary tools. In terms of Vincent videos, users can directly input text prompt words to generate videos, support movie-level lens motion design (such as wide-angle lenses, low-angle zoom-in, etc.), have strong semantic understanding capabilities, and support the highest output resolution of 1080p high-definition. The generation speed is officially claimed to be 10 seconds for a video. Users reported that it takes tens of seconds to generate a 480p video, which is suitable for efficient short video creation scenarios. When it comes to Tusheng videos, Vidu offers two exclusive features. The first and last frame control function allows users to upload starting and ending frame images, and AI automatically generates a smooth transition animation between the two frames. The animation function can transform static comics or illustration images into natural and smooth dynamic videos. In addition, it also supports customizing motion effects through text prompt words. Reference video is the core highlight feature of Vidu and it is also the world's first reference video generation technology. This feature supports multi-subject consistency, uploading up to 7 reference images of people, objects or scenes to ensure that these elements are consistent in the generated video. The main element library function allows users to save reusable characters, props and scenes into their personal library, enabling one-click selection and improving creative efficiency. The feature also supports blending 2 to 3 more reference visual elements into the same seamless video. In terms of auxiliary tools, Vidu provides a template library, including popular short video templates such as AI kissing, hugging, object growth, AI dress-up and other trendy special effects. Vidu Claw provides related capabilities as a functional module. The API open platform provides business integration interfaces for enterprise-level users and supports batch calls. In terms of user experience, Vidu's multi-subject consistency performance is industry-leading, and there will be no drift problems in characters, objects or scenes in the generated videos. Character movements are natural and non-mechanical, and it excels in 2D animation, anime content and manga adaptations. The dynamic range is large and supports rich multi-angle action sequences. The low barrier to entry requires no advanced video production skills and supports flexible output from low resolution to 1080p. All content uploaded by users is strictly confidential, and privacy protection has high priority.

  • Vidu adopts a points-based pricing model. Registered users can receive a one-time free point allocation and can start video generation without paying. Offers unlimited free credits and unlimited video generation capabilities during off-peak hours. Paid subscription packages are divided into different levels. The specific prices are not displayed on the homepage of the official website. You need to visit the pricing page for details. Enterprise users can obtain commercial integration capabilities through the API open platform (https://platform.vidu.cn), which is billed based on the amount of calls.

  • In terms of positive reviews, users recognize Vidu’s technical strength as a domestic self-developed video generation model. The speed of generating a video in 10 seconds greatly improves the creation efficiency. The multi-subject consistency of the Reference Video feature is a core competency, with characters and objects remaining stable across successive frames. The animation effect is natural and smooth, which is very suitable for two-dimensional content creators. The dynamic range is large, the generated video is rich in action, and it supports movie-level lens movement design. Unlimited free mode during off-peak periods is light user friendly. Data security is guaranteed, and content uploaded by users is strictly confidential. In terms of negative feedback, the price of the paid package is too high for individual users, and there are certain restrictions when using it completely free of charge. During peak periods, you need to queue up and wait for generation, and video generation requires a certain time cost. The functional learning curve is steep, and newbies need time to become familiar with various functions. There is still room for improvement in the understanding and generation of some complex scenes.

  • From an industry perspective, Vidu is the first video generation model of the same level in China after the release of OpenAI Sora, and is rated as "the first echelon in the world" by professional users. In February 2026, Vidu Q3 ranked first in China and second in the world in the latest list released by Artificial Analysis, an authoritative international AI evaluation agency, surpassing competing products such as Runway Gen-4.5 and Google Veo 3. In April 2026, Vidu Q3 ranked first with high scores in the special evaluation list SuperCLUE-ComicShorts, proving its leading position in the field of reference student videos. Shengshu Technology has received investment from leading companies such as Alibaba, providing financial guarantee for continued product iteration. In terms of competitive product landscape, Vidu faces dual competition from both international and domestic sources. Internationally, products such as OpenAI Sora, Google Veo 3, and Runway Gen-4 have strong technical capabilities. Domestically, competing products such as Keling (owned by Kuaishou) and Wanxiang (owned by ByteDance) are also rapidly iterating. Vidu's differentiated advantage lies in its reference video function and multi-subject consistency technology, but it still needs to be continuously optimized in terms of generation speed and cost.

  • The main disputes and risks include: As a commercial product, the payment model may limit the willingness of some users to use it. The quality of video generation is still unstable, and complex scenes are sometimes difficult to accurately represent. There are legal ambiguities surrounding the copyright ownership and commercial use scope of AI-generated content. The market competition is fierce, technology iteration is fast, and there is a risk of being surpassed by competitors.

  • Vidu is suitable for the following people: First, social media content creators who need to quickly produce eye-catching short videos. The second is a professional film and television animation studio that needs assistance in generating movie-level animation and action sequence supplements. The third is the commercial marketing team, which needs marketing video optimization such as product image animation and advertising background replacement. The fourth is creators who have high requirements for video quality and are not satisfied with simple templated content. It is not suitable for the following groups: First, users who are completely pursuing free use. Although there is a free quota, the functions are limited. The second is scenes that need to be generated in real time, and you need to wait during peak periods. Third, for users who are extremely cost-sensitive, batch production costs are high. In terms of usage suggestions, you can make full use of the unlimited free mode during off-peak periods for practice and creation. The reference video function is the core advantage. It is recommended to study in depth how to effectively use reference pictures. Follow the official community and WeChat group to get the latest feature updates and usage tips. Used in conjunction with other video editing tools, Vidu is responsible for generating the first draft and leaving the fine adjustments to professional tools later.

  • As the first purely self-developed AI video generation platform in China, Vidu AI has reached the world's first-class level in terms of technical strength. The reference video function is the core competitiveness, and the multi-subject consistency technology leads the world. The Q3 version ranked first in China and second in the world in international evaluations, proving that its product strength has been widely recognized. For domestic users, Vidu provides high-quality video generation services that can be used without circumventing the wall, and supports Chinese prompt words, which is more in line with domestic user habits. As technology continues to iterate and the ecosystem gradually improves, Vidu is expected to become the preferred platform for Chinese AI video creation. However, it should be noted that its payment model has certain thresholds for individual users, and peak generation requires waiting. It is recommended to plan usage time and budget reasonably.

User Reviews

  • 头像
    UGRLYR
    After using Vidu for two weeks, I have to say that the reference student video function is really powerful! You can maintain consistency by uploading several character images, and generate continuous images without any face blindness.

  • 头像
    PEwol
    As a self-developed AI video tool in China, Vidu is indeed up to the mark! Generating a video in 10 seconds is not a boast, it is very efficient.

  • 头像
    Shirley.Cox
    The unlimited free mode during non-peak periods is so delicious! There is no need to queue up for liver videos at night, and unlimited generation is possible.

  • 头像
    Logan_Smith_2021
    The first and last frame control function of Tusheng Video is my favorite. Just two key frames can generate a super smooth transition animation!

  • 头像
    5ldqj_0xdz
    The animation effect is really amazing to me. Static comics go in, and vivid videos come out. This is good news for two-dimensional creators.

  • 头像
    COtur
    Multi-agent consistency is the most stable one I have ever used! Using 7 reference pictures together, the characters and props will not drift.

  • 头像
    JRobinson_66096
    The main element library function is so practical. Characters and scenes can be saved and directly called with one click next time, doubling the creation efficiency.

  • 头像
    Kelly.Gonzales_Max
    You still have to wait in line for a while during peak periods. It's not as fast as some platforms, but the quality is worth the wait.

  • 头像
    常敏
    The paid package is a bit expensive for students, and I hope there will be a more affordable personal version.

  • 头像
    LambdaLP549
    It is comparable to products from major international manufacturers, and it supports Chinese prompt words, making it easier to use.

  • 头像
    PhilipJohnson_2020
    The Q3 version has been upgraded with stronger functions, including six special effects, five sound effects and four scenes, giving it a larger creative space.

  • 头像
    Jason_Morales520
    The movie-level lens movement design is so professional, and the videos taken from wide angles and low angles have a blockbuster feel.

  • 头像
    TokqnTigerBell
    It does a good job in terms of data security. All uploaded content will be kept strictly confidential, so you can use it with confidence.

  • 头像
    JesseSimmonsZ
    The difficulty of getting started is much lower than Runway, and no professional knowledge is required to get started.

  • 头像
    CarlEdwards
    Like the international evaluation rankings, the generation effect can indeed reach the level of the second tier in the world.

  • 头像
    DYrus
    The popular special effects in the template library are very interesting. AI kissing and dressing up are very popular and are suitable for making short videos.

  • 头像
    AyselLamm
    Sometimes the generation of complex scenes will still have flaws, which requires several more attempts or post-production adjustments.

  • 头像
    Ruth509
    The reference video function is the most powerful one I have ever used, and the integration of multiple elements is stress-free.

  • 头像
    RuthButler520
    Compared with competing products such as Kelingwanxiang, Vidu's multi-agent consistency is the most stable.