In-depth Report
-
Vidu S1 is a real-time interactive video basic model officially released by Shengshu Technology at the 2026 Global Digital Economy Conference (July 3). It advances AI video from "offline generation of a complete video" to "continuous online, chatting and generation" real-time interactive form. Its core indicators are 540P (960×540), steady-state 25FPS, and peak 42FPS. The end-to-end delay is generally reduced to less than 3 seconds. It supports real-time voice control of character actions, instant creation of digital people in a single picture, and unlimited continuous interaction. So far, it has opened small-scale online internal testing and API commercial access, mainly for scenarios such as emotional companionship, virtual idols, interactive live broadcasts, online education, and in-car assistants.
-
Shengshu Technology was founded by Professor Zhu Jun of Tsinghua University. It is one of the earliest AI video companies in China that focuses on the "universal world model" route. Its Vidu series has previously established a firm foothold in the creator circle with its functions such as graphic video, multi-subject consistency, and reference video. Vidu S1 was developed by Zhang Jintao, a post-2000s doctoral student led by Zhu Jun, as the R&D director to complete the full-link development, positioning it as a key step in the direction of "real-time interactive generation" of the general world model. On the day of its release on July 3, 2026, Shengshu Technology was also selected as a benchmark enterprise for new models and new applications in the "2025 Beijing Digital Economy Benchmark Enterprises" of the Beijing Software and Information Services Industry Association. At the same time, a leading car company also announced that its 2026 flagship model will be equipped with a customized version of Vidu S1 for the first time.
-
The biggest change in Vidu S1 is the interaction paradigm. The traditional large video model is an offline process of "submit request → wait for generation → playback result", while S1 was changed to "real-time voice input → frame-by-frame generation → continuous playback → continuous response". The model is already calculating subsequent frames while playing the current frame, and new voice commands can be instantly injected into subsequent generations, allowing the picture to change in real time with the conversation. When it comes to character creation, S1 cuts the threshold to an extremely low level: users only need to upload a reference picture (real person, animation, cute pet, or game character), and the model will understand the identity, appearance, and visual style on the spot, and generate synchronized mouth shapes, expressions, eyes, gestures, and body movements in real time, without the need for modeling, binding, mouth shape adaptation, or character-specific training. At the sound level, it supports system timbre or recording your own voice for customization, ensuring that the visual identity is consistent with the vocal identity. In actual testing, voice control behavior is the ability that attracts the most attention and has the greatest gap. Simple commands such as "raise your left hand", "push your glasses" and "tick your hair" can mostly be completed smoothly. Compound commands (such as arranging your hair with your left hand, arranging your clothes with your right hand, and comparing your heart with your hands) can be done to a minimum but not enough. However, continuous large movements such as turning, dancing, and squatting seem to be restricted at the system level, and the character will only shake slightly. Users ridiculed that "the mouth promises to be positive, and the body only shakes twice symbolically." Delay exists objectively - after finishing a sentence, the character often has to pause for a moment before responding. It is fast when used in video generation, but it becomes obvious when put into a video call. Moreover, the character staring at you all the time can easily make people feel oppressive.
-
S1's real-time interaction is charged based on the generation time. Public information shows that API billing is approximately 3 points/2 seconds, which is deducted every 6 seconds and rounded up according to a 2-second cycle. New users are given 1,000 free trial points (approximately 11 minutes of real-time interaction), and can experience the complete API, all sounds and languages. The unit price of points is about 0.03125, which is equivalent to about 2.8 yuan/minute, which is about 168 yuan per hour. There is also a daily basic free interaction time during the internal testing phase (with watermark, limited frame rate, and not for commercial use). The enterprise side provides exclusive SLA packages, which can purchase high concurrency, watermark removal, 42FPS full permissions and customized character/voice cloning access. There is currently no fixed monthly membership, and points are consumed based on the actual interaction time.
-
Positive voices focus on "zero threshold" and "like real people". Many experiencers believe that the three points of creating a character with a single picture, voice-driven full-dimensional behavior, and unlimited continuous dialogue are generational advantages. Chatting all night will not cause the status to fade away, the mouth shape can be kept up, and the memory can continue the previous dialogue. Being an AI companion and a virtual idol has a "living feeling". The digital person who built the "cold-hearted academic master" who was a tester was good at giving lectures, had warm interactions, and could remember what was discussed previously and summarize it later. Negative feedback focused on realism and latency. Actual tests by multiple media outlets have pointed out that complex movements can be confusing, character descriptions are disconnected from actual behaviors, and the recording of one's own voice occasionally fails to work and the voice lines are cut randomly. The half-duplex experience (similar to a walkie-talkie, you say something to it and it will not shut up when you interrupt it) has been compared with pure voice products such as Doubao. It is believed that the 10-second delay is too modern in the "human-machine love companionship" scenario. The title of an experience article by NetEase Lei Technology directly states "There are more problems than expected." The roll call delay is large, the movements are drawn to death, the body shape and movements do not match, and the teeth and cheeks are occasionally inconsistent when the expression is excited.
-
The industry generally regards S1 as a landmark product that has shifted the video generation track from "competition on production quality" to "competition on real-time interaction". Compared with competing products such as HeyGen, D-ID, and Synthesia that are good at offline digital people, S1 has generational advantages in depth of interaction and sustained capabilities, and is especially suitable for companionship, live broadcasts, and NPC scenes that require a "real-life feel." However, it still has shortcomings in 4K film and television-level image quality and precise facial feature adjustment. Offline feature film production is not as good as the higher resolution and fine editing of HeyGen/Synthesia. On the technical side, it relies on Shengshu's self-developed TurboDiffusion inference acceleration framework and TurboServe streaming deployment engine, and superimposes attention optimization and model quantification such as SageAttention, SLA, SpargeAttention, etc., to reduce the delay of real-time video generation to the passing line on consumer-grade graphics cards. Some comments juxtapose it with MaineCoon Chat Mode of the same period, believing that the overall real-time video digital human product is still in the "pre-modern" stage, but the direction is clear.
-
The main controversy focuses on the management of experience expectations: there is a gap between the officially promoted "semantic perception, intention understanding, and corresponding comfort response" and the actual test. Voice control behavior is unstable on complex instructions, which can easily make users feel "fooled." The somatosensory shortcomings of half-duplex and higher latency in companion scenes have also triggered discussions on "whether virtual companionship is really ready." In addition, 540P resolution is still a hard limit for e-commerce live broadcasts and film and television promotions that require high-definition appearances; character details are not as controllable as 3D modeling solutions, which also means that it has temporarily given way to high-precision professional production.
-
It is suitable for: individual creators who want to be exclusive digital people at low cost, live broadcast teams who need to be online 24 hours a day, game developers who want to interact with NPCs in real time, companies who need brand digital people for customer service, and teams who want to quickly verify real-time companionship/education products. Not suitable for: professional film and television productions that pursue 4K film and television quality, projects that require complex multi-character scenes or 3D high-precision modeling. In terms of usage recommendations, individual creators should first use free credits to test the matching of actions and sounds before deciding whether to use it commercially; it is more stable for commercial teams such as e-commerce and education to directly use the enterprise API package. In terms of alternatives, low-latency products such as Doubao are available for pure voice companionship, and HeyGen and Synthesia are available for offline digital human short film production.
-
Vidu S1 transforms AI video from "shooting a short video" to "opening a real-time digital human call". Zero training on a single picture, voice-driven full-dimensional behavior, and unlimited time interaction do constitute a generational difference. However, delay, action controllability, and realism are still hurdles it must overcome before it can be accompanied by the public.
User Reviews
-
BillyBaker_2023—单图就能捏数字人,门槛真的低。 -
龙昊—540P 画质确实糊,拿它当视频通话还行,当正经出片就别想了。 -
钱怡凤—简单指令「举手」很顺,复合指令就勉强,复杂动作直接糊弄。 -
郭丹丹—体验下来延迟是真不小,你说完一句它还要愣一下才开口。放进视频生成算快,可一旦当成视频通话,这种停顿就特别出戏。 -
傅彤晨—我录了自己的声音想做定制音色,结果完全没生效,画面里随机切男女声,看着相当诡异,这点希望能修。 -
Patrick.Gray—让它「我心情不好」它会做安抚动作,理论上感知情绪挺酷,但实际跟官方宣传的理想状态还有落差,别指望太智能。 -
Virginia.Bell00702—算了一笔账有点劝退:实时交互按 2 秒扣 3 积分,折算差不多 2.8 元一分钟,一小时就是一百六。我前后砸了两千多积分、一千多秒连线才把整套多模态闭环评测跑完,个人玩玩免费额度够了,但真要拿来做电商直播或者陪伴,长尾成本得先算清楚,不然几分钟就烧没了。建议商用团队直接走企业 API 套餐,别拿免费额度硬扛。 -
AXmoo—用自动化脚本接了 API 跑压测,预设角色库从最早 6 个扩到 11 个,迭代速度挺猛。两组长时对话下来画面基本钉在 25 帧、零抖动,聊久了脸也没漂,工程稳定性确实比想象中稳。我用 Claude 一个多小时就跑通了接入,连哈兰德和沃齐尼亚都拉来做了中英文压测。就是接入要啃文档加阿里云 ARTC SDK,对纯小白来说门槛不低,但愿意折腾的程序员基本一天能上手。 -
JRogersZ—消费级显卡就能跑 540P 实时生成这点很香,小团队和个人创作者不用堆 A100 也能落地,本地部署门槛被砍了一大截。 -
ticklishgorilla525—跟 HeyGen、D-ID 比的话,S1 在交互深度和持续对话上确实是代际领先,但 540P 分辨率摆在那,正经影视级制作还是得靠离线方案,定位要分清楚。