In-depth Report
-
On July 20, 2026, HiDream.ai released the world's first unlimited content creation multi-modal agent vivago R1 at the World Artificial Intelligence Conference WAIC. This product completely rewrites the inherent model of "15-second short shots" in the AI video generation industry, using a multi-agent collaboration architecture to achieve full-link automation from narrative conception to finished film distribution. The company completed three consecutive rounds of financing within three months, with a total of more than 2.1 billion yuan, led by national capitals such as the Social Security Fund, and its valuation has ranked among the unicorns. The HiDream-O1-Image-1.5 model ranked second in the world and first in China with 1265 ELO in the Artificial Analysis evaluation.
-
HiDream.ai was founded in March 2023 and is headquartered in Hefei High-tech Zone. It is a generative AI company in the field of visual multi-modal basic model development and application. The founder, Dr. Mei Tao, is a Ph.D. in the Department of Automation at the University of Science and Technology of China and a foreign academician of the Canadian Academy of Engineering. He served as a senior researcher at Microsoft Research Asia for 12 years and later served as the vice president of JD.com, responsible for the JD AI Research Institute and Multimedia Laboratory. The company's technical route has chosen a "hard core" since its inception - the native all-modal architecture UiT (Unified Transformer), abandoning the traditional "text encoder + VAE + diffusion model" modular path, mapping original signals such as image pixels, text tokens, video voxels, etc. into the same shared token space, and a unified set of Transformers to complete understanding, generation and reasoning. 2026 is the year of the explosion of the future of smart elephants. In just three months, the company completed three consecutive rounds of financing: more than 500 million yuan in April (led by Dongfang Fuhai, Anhui Investment Group, etc.), a new round of capital injection in May (Shenzhen Venture Capital, Jinpu Investment, etc.), and a C round of 1.5 billion yuan in July, jointly led by the social security fund Sichuan Revitalization Science and Technology Innovation Fund, ICBC Capital, etc., with a total of more than 2.1 billion yuan. Investors include the national team, local state-owned assets, film and television industry capital and top financial institutions. The "luxury" lineup of this round of financing is extremely rare among large AI model companies.
-
The core concept of vivago R1 can be summarized in one sentence: it is not a black box with "cue words in, video out", but a multi-modal intelligent agent that can think like a director. Users only need to upload a few photos and short text descriptions to generate a minute-level continuous narrative video. Three-layer multi-agent architecture The core design of vivago R1 is hierarchical memory and multi-Agent collaboration architecture. The top layer is the orchestrator. Like a general director, it only precipitates the final draft - character settings, narrative main lines, confirmed decisions, and filters out process noise. In the middle is the Specialist Agents, which are stationed with Screenwriter Agents, Director Agents, Designer Agents, and Editor Agents. Each has its own independent memory and is responsible for its own professional links. The bottom is the execution layer (Execution Agents), which is responsible for segmented and concurrent generation. Each sub-Agent only processes a segment of the task and leaves immediately. If a segment fails, only that segment will be retried, without affecting the overall progress. Mei Tao summarized the operating logic of this mechanism as "the memory becomes lighter layer by layer, and the final version floats layer by layer" - which is very consistent with the intuition of real creators. SkillHub creative method library vivago R1 turns the creative method into a product. Skill is no longer a simple function button, but compresses professional processes such as script disassembly, lens organization, material inheritance, and style control into "methods" that users can directly call. Users only need to know whether they want to do product advertising, short stories, or brand content, and Skill is responsible for translating the goals into "work agreements" that Agent can execute. Up to now, SkillHub has accumulated hundreds of creative skills that have been verified by real projects, covering cross-border e-commerce materials, film and television storyboards, plot account operations and other scenarios. Dual view interaction The product provides both Chat and Canvas views. Chat is a task advancement view, responsible for undertaking user intentions, and is the "timeline" for continuous collaboration between users and Agents; Canvas is a project context view that reorganizes key content in the creative process into a browsable and reusable result space. After the video is completed, it can be directly bound to TikTok, YouTube and other overseas accounts for one-click distribution. Tested performanceAccording to TMTpost's actual measurement, if you enter the command "Create a 1-minute continuous narrative life short film with the theme of an ordinary day of ordinary people", the system will automatically call the multi-lens narrative process, split the complete story into multiple scenes, and activate multi-Agent division of labor. The generated results show excellent character consistency - no matter under natural light, indoor cold light or night street lights, the facial features, hairstyle and body posture of the characters always remain highly unified, the picture quality reaches the film-level level, the transition of light and shadow is natural, and the movements are coherent without lag.
-
vivago adopts a "free trial + tiered payment" model. Overseas version pricing: free version (daily free points, 5-second video with watermark), basic version ($7.9/month, 1,000 points, no watermark), enhanced version ($19.9/month, 3,800 points, 4K image quality), professional version ($59.9/month, 12,000 points, AI Chat Agent). The domestic version is more affordable, with the personal paid version priced at 29 yuan/month and the enterprise version priced at 299 yuan/month. Zhixiang Future has built a "1+1+3" business model: 1 HiDream-O1 series large model base + 1 Token Hub platform + 3 major scenarios (commercial marketing, film and television creation, content creation). The company's products have covered more than 50 million users and more than 40,000 corporate customers in more than 100 countries and regions around the world.
-
Overseas content creators have given positive feedback to vivago. In May 2026, the grayscale version topped the daily list on Product Hunt. In professional reviews, Titanium Media gave it a 90-point rating. Users generally recognize its consistent performance in long videos, especially its character stability and picture quality, which are far superior to similar products. Negative feedback mainly focuses on: the free version has an obvious watermark, the cost of using advanced features under the points system is higher, it still takes a long time to generate long videos, and there are still occasional problems with character consistency in some complex scenes. In the evaluation of domestic users, its features such as being able to access without a special network, supporting mobile phone number/WeChat registration, Chinese interface and customer service have achieved high satisfaction.
-
Many industry observers believe that vivago R1 marks a paradigm shift in AI video generation from "single-point tool" to "full-link creation system". Its underlying UiT native all-modal architecture is seen as the critical path to the world model. At the technical level, HiDream-O1-Image-1.5's performance on the Artificial Analysis Vincent chart list (1265 ELO, second in the world, surpassing Google Nano Banana 2, NVIDIA Cosmos3 and ByteDance Seedream 4.0) has established HiDream-O1-Image-1.5's first-tier status in the field of visual generation. From a capital perspective, the entry of social security funds is interpreted as an important layout of national patient capital in the direction of the original full-modal world model. The joint emergence of film and television industry capital (Huace Film and Television, Shanghai Film New Vision Fund, Hubei Yangtze River Film Group) shows the trend of deep integration of AI and the film and television industry.
-
The field of AI video generation still faces copyright ownership disputes. Although Zhixiang clearly understands that users have complete copyright on the generated content, the compliance and potential infringement risks of training data are still industry-wide problems. In addition, the risk of "deep forgery" in AI-generated videos also requires continued attention - if high-quality AI videos of unlimited duration are used to spread false information, they may bring social risks. From the perspective of market competition, competing products such as OpenAI's Sora, Google's Veo, and ByteDance's Jimeng continue to develop. Whether Zhixiang can maintain its first-mover advantage in the future remains to be seen. The high cost of model research and development (2.1 billion in financing in three months) also places higher demands on the pace of commercialization.
-
vivago R1 is most suitable for the following groups: social media creators who need to produce short video content at a high frequency, e-commerce merchants who lack a professional video production team, and small and medium-sized enterprises who need to quickly produce brand promotion materials. Not suitable for professional film and television producers who require precise control of a single shot (they may need more fine-grained control tools such as Pika). For domestic users, it is recommended to start with the free version to experience it, and then upgrade to the paid version after confirming that it is suitable. Enterprise users can directly choose the enterprise version or customized version to obtain 4K image quality and copyright certification.
-
vivago R1 is one of the most breakthrough products in 2026 in the field of AI video generation. Its architectural innovation with multi-agent collaboration as its core has successfully upgraded AI video from a "material generator" to a "creation agent", allowing ordinary users to complete long video creation with one click that previously required a professional team. As UiT's native full-modal architecture continues to iterate, Intelligent Vision is moving towards a broader AI infrastructure along the path of "image → video → space → action → world model" in the future. For content creators and business users, there’s never been a better time to get on board.
User Reviews
-
Web3DevLarsen—作为一个跨境电商卖家,我现在每天用vivago批量生成产品展示视频。它的SkillHub里的电商模板特别好用,选好品类后系统自动匹配视觉风格和视频节奏,生成的视频完播率比我自己剪的高了将近30%。唯一的问题是免费版每天积分不够用,但我已升级Pro了,算下来比自己招剪辑师便宜太多了。 -
Patrick.RiveraQ—昨天用 vivago R1 做了个1分钟的剧情短片,真的不用学任何剪辑软件,输入一段话它就自动帮你拆分成好几个镜头。最让我惊讶的是角色从头到尾都没变脸,之前用其他AI视频工具最烦的就是换脸问题。 -
Jerry_PetersonIII—刚升级了Pro版台服会员,59.9刀一个月说实话有点肉疼,但生成长视频的质量确实比其他工具好太多了。做跨境电商的素材以前要外包给后期团队,现在自己一个人就能搞定。 -
i6prb—我在一个小型广告公司做创意总监,上个月开始团队全面用vivago R1出短视频素材。说实话一开始我是抵触的,觉得AI做的东西没有灵魂。但用了两周之后我完全改观了——不是AI做得好,而是客户不在乎。客户只看两件事:能不能准时交付和画面够不够吸引人。以前一个brief下来我们要两三天出demo,现在半天就能给客户看初版,客户确认后直接批量出片。这已经不是好不好用的问题了,而是整个行业的游戏规则被改变了。 -
ScottScott—太强了!!直接绑定TikTok一键分发是我最看重的功能。以前做一条短视频要导出→上传→调参数→写标题→发布,现在一个平台全搞定,效率翻倍。 -
菊花_18—免费版积分太少,宣传片中能自由编辑倒是好用,就是积分制不够透明,都打水印有点烦人,再观望一段时间。 -
冯贞轩—试了一下免费的每天可以生成几次,5秒视频带水印,当个尝鲜体验还行。但真要商用还是得付费版,这价格对个人创作者来说门槛还是有点高啊。 -
Nancy_GarciaX—我是做影视后期的,说实话vivago R1生成的效果确实惊艳,但离专业制作还有距离。不过对于短视频创作者和省钱的中小商家来说绝对够用了。 -
Denise_Ramos_99—等了这么久终于发布了!WAIC上看到演示的时候就被震撼到了,这就是我一直想要的AI视频工具,不再是15秒片段而是真正有叙事的长视频。 -
smallgoose248—WAIC上看了vivago R1的现场演示,真的被震撼到了。它把整个视频创作流程从构思、分镜、素材生成到成片输出全部自动化了。以前做一个3分钟的剧情短片,至少需要一个编剧、一个导演、一个剪辑师配合好几天才能完成。现在一个人一台电脑,输入一段描述就能搞定。这不是效率提升,这是生产关系的变化。 -
joKOC—作为短视频运营我直接无脑冲了。日更3条的压力真的大,vivago能帮我批量出素材,角色一致性好这个点太关键了,以前用Stable Diffusion生成的角色每张图都像换了一个人,根本没法用于品牌内容。这次看到R1的演示,角色在多个镜头和不同光线环境下都能保持统一,我终于可以放心用AI做品牌视频了。我算了一下,每个月花60刀省了一个剪辑师的工资,这个ROI怎么算都划算。 -
ERussellII—用了两周vivago,感觉这玩意儿就是TikTok卖货的神器。上传几张产品图,写一段描述,自动帮你生成带货视频,连话术都配好了。 -
蝴蝶505—中文渲染比Sora和Runway好太多了。我之前用英文工具生成中文字幕,十个字里有八个是错的。vivago R1至少中文能看得懂。 -
SIcoo—我研究了一下它的技术架构,底层用的是自研的UiT原生全模态架构,把文本、图像、视频、音频都放在同一个空间里处理,不需要像传统方案那样在多个模型之间来回翻译信息。这就是为什么它能保持角色一致性那么好的根本原因。虽然不太懂技术细节,但从生成效果来看这条路确实走对了。 -
TEgon—听说公司三个月融了21个亿,资方还有社保基金和工银资本,这种国家队背景至少不用担心跑路问题了。融了这么多钱应该能支撑很长时间的研发,不像有些AI公司融完A轮就撑不住了。而且从他们的技术路线来看,UiT原生全模态这个方向确实比市面上其他方案靠谱,希望他们能坚持走下去。 -
RitaPrice—写了个脚本让它生成一个3分钟短剧,竟然一遍就过了,角色没崩、场景切换流畅、配乐也都卡在节奏上。这种「一遍过」的体验在其他AI视频工具上从来没遇到过,我真的枯了。 -
JessicaCook_2024—测试了Hidream的图像编辑功能,局部修改的一致性确实比同期的其他工具好。不过修改后画质会有轻微损失,大概压缩了15%的样子,希望后续能优化。 -
Doris.BrownZ—我觉得vivago R1最大的价值不是生成视频本身,而是它把整个创作流程串起来了。从构思到分镜到成片到分发,一个平台搞定,这才是真正的生产力工具。 -
CodyMartinez—上周用vivago给一个美妆客户做了一批TikTok素材,客户要求出20条不同角度的产品展示视频。我用vivago的批量生成功能,上传了产品图库和卖点文案,系统自动拆分成不同的脚本和镜头组合,半天就出了初版。客户对角色一致性和画面质感都很满意,说比他们之前找的MCN机构做得好。这个工具对中小代理公司来说真的是降维打击。 -
KAram—用vivago做了个品牌宣传片,3分钟时长,客户反馈说画面质感出乎意料的好。以前这种项目至少要花一两万找团队拍,现在自己用AI几分钟就能出个初版。 -
贾博萱—生成速度有点慢,做3分钟视频要等5到10分钟。不过想想以前一个镜头要用After Effects渲染几个小时,这个速度也算能接受了。 -
JoseRossK301—SkillHub里的模板是真的多,跨境电商、剧情号、产品广告什么场景都有。关键是不用自己调参数,选好模板填内容就行,对新手特别友好。 -
JerryRivera_Plus—角色一致性确实是一大亮点。我之前测试过其他AI视频工具,比如Runway和Pika,生成人物稍微换个角度脸就不一样了,根本没法用在需要固定角色的剧情内容上。vivago在这方面做得是最稳的,它底层用的UiT架构把文本、图像、视频都统一建模了,理论上就比其他拼接式的方案好。不过偶尔在复杂场景下人物还是会有轻微变化,还有优化空间。 -
crazylion152—国内版29元一个月还算良心,比海外版便宜不少。而且直接手机号就能注册,不用折腾海外账号,这点加分。 -
TheSimonaKishka_x—用了vivago R1两个月了,从灰度版一路跟到正式版,说一下我的真实感受。优点是角色一致性、中文理解能力和长视频生成这三个维度确实吊打同行。缺点也很明显:一是收费偏贵,Pro档60刀一个月对个人来说负担不小;二是输入复杂场景描述时经常有理解偏差,需要反复调prompt;三是敏感词检测有时候莫名其妙,我写「保健品」三个字就违规了,搞不懂。 -
Helen.Stephens_99—免费用的很开心,看完才能彻底爱上,深度体验后想一直用,建议大家也试试。目前还没付费,但已经在考虑买plus了。 -
move8lcn—想请教一下各位,vivago R1生成的长视频有没有版权问题?能不能直接用在做商业广告?官方说plus以上有商业授权,但我不太确定条款细节。 -
EMurrayQ998—HiDream-O1-Image-1.5那个生图模型真的强,1265分全球第二,我用它做了几张产品图完全看不出是AI生成的。相比之下,vivago的底层能力确实有保障。 -
邹梦—敏感词检测有点迷,我一开始没弄好,总提示违规。反复修改了几次总算能正常运行了。 -
Jordan.Cook007012—海外博主圈最近都在推vivago,我关注的一个YouTuber专门做了两期评测视频。能感受到海外用户对无限时长这个特性真的非常兴奋,毕竟以前只能做15秒的片段太局限了。