Vidu S1
The real-time interactive video basic model released by Shengshu Technology supports real-time voice control of characters, single-image creation of digital people, and unlimited duration of continuous interaction.
In-depth Report
-
Vidu S1 is a real-time interactive video basic model officially released by Shengshu Technology at the 2026 Global Digital Economy Conference (July 3). It advances AI video from "offline generation of a complete video" to "continuous online, chatting and generation" real-time interactive form. Its core indicators are 540P (960×540), steady-state 25FPS, and peak 42FPS. The end-to-end delay is generally reduced to less than 3 seconds. It supports real-time voice control of character actions, instant creation of digital people in a single picture, and unlimited continuous interaction. So far, it has opened small-scale online internal testing and API commercial access, mainly for scenarios such as emotional companionship, virtual idols, interactive live broadcasts, online education, and in-car assistants.
-
Shengshu Technology was founded by Professor Zhu Jun of Tsinghua University. It is one of the earliest AI video companies in China that focuses on the "universal world model" route. Its Vidu series has previously established a firm foothold in the creator circle with its functions such as graphic video, multi-subject consistency, and reference video. Vidu S1 was developed by Zhang Jintao, a post-2000s doctoral student led by Zhu Jun, as the R&D director to complete the full-link development, positioning it as a key step in the direction of "real-time interactive generation" of the general world model. On the day of its release on July 3, 2026, Shengshu Technology was also selected as a benchmark enterprise for new models and new applications in the "2025 Beijing Digital Economy Benchmark Enterprises" of the Beijing Software and Information Services Industry Association. At the same time, a leading car company also announced that its 2026 flagship model will be equipped with a customized version of Vidu S1 for the first time.
-
The biggest change in Vidu S1 is the interaction paradigm. The traditional large video model is an offline process of "submit request → wait for generation → playback result", while S1 was changed to "real-time voice input → frame-by-frame generation → continuous playback → continuous response". The model is already calculating subsequent frames while playing the current frame, and new voice commands can be instantly injected into subsequent generations, allowing the picture to change in real time with the conversation. When it comes to character creation, S1 cuts the threshold to an extremely low level: users only need to upload a reference picture (real person, animation, cute pet, or game character), and the model will understand the identity, appearance, and visual style on the spot, and generate synchronized mouth shapes, expressions, eyes, gestures, and body movements in real time, without the need for modeling, binding, mouth shape adaptation, or character-specific training. At the sound level, it supports system timbre or recording your own voice for customization, ensuring that the visual identity is consistent with the vocal identity. In actual testing, voice control behavior is the ability that attracts the most attention and has the greatest gap. Simple commands such as "raise your left hand", "push your glasses" and "tick your hair" can mostly be completed smoothly. Compound commands (such as arranging your hair with your left hand, arranging your clothes with your right hand, and comparing your heart with your hands) can be done to a minimum but not enough. However, continuous large movements such as turning, dancing, and squatting seem to be restricted at the system level, and the character will only shake slightly. Users ridiculed that "the mouth promises to be positive, and the body only shakes twice symbolically." Delay exists objectively - after finishing a sentence, the character often has to pause for a moment before responding. It is fast when used in video generation, but it becomes obvious when put into a video call. Moreover, the character staring at you all the time can easily make people feel oppressive.
-
S1's real-time interaction is charged based on the generation time. Public information shows that API billing is approximately 3 points/2 seconds, which is deducted every 6 seconds and rounded up according to a 2-second cycle. New users are given 1,000 free trial points (approximately 11 minutes of real-time interaction), and can experience the complete API, all sounds and languages. The unit price of points is about 0.03125, which is equivalent to about 2.8 yuan/minute, which is about 168 yuan per hour. There is also a daily basic free interaction time during the internal testing phase (with watermark, limited frame rate, and not for commercial use). The enterprise side provides exclusive SLA packages, which can purchase high concurrency, watermark removal, 42FPS full permissions and customized character/voice cloning access. There is currently no fixed monthly membership, and points are consumed based on the actual interaction time.
-
Positive voices focus on "zero threshold" and "like real people". Many experiencers believe that the three points of creating a character with a single picture, voice-driven full-dimensional behavior, and unlimited continuous dialogue are generational advantages. Chatting all night will not cause the status to fade away, the mouth shape can be kept up, and the memory can continue the previous dialogue. Being an AI companion and a virtual idol has a "living feeling". The digital person who built the "cold-hearted academic master" who was a tester was good at giving lectures, had warm interactions, and could remember what was discussed previously and summarize it later. Negative feedback focused on realism and latency. Actual tests by multiple media outlets have pointed out that complex movements can be confusing, character descriptions are disconnected from actual behaviors, and the recording of one's own voice occasionally fails to work and the voice lines are cut randomly. The half-duplex experience (similar to a walkie-talkie, you say something to it and it will not shut up when you interrupt it) has been compared with pure voice products such as Doubao. It is believed that the 10-second delay is too modern in the "human-machine love companionship" scenario. The title of an experience article by NetEase Lei Technology directly states "There are more problems than expected." The roll call delay is large, the movements are drawn to death, the body shape and movements do not match, and the teeth and cheeks are occasionally inconsistent when the expression is excited.
-
The industry generally regards S1 as a landmark product that has shifted the video generation track from "competition on production quality" to "competition on real-time interaction". Compared with competing products such as HeyGen, D-ID, and Synthesia that are good at offline digital people, S1 has generational advantages in depth of interaction and sustained capabilities, and is especially suitable for companionship, live broadcasts, and NPC scenes that require a "real-life feel." However, it still has shortcomings in 4K film and television-level image quality and precise facial feature adjustment. Offline feature film production is not as good as the higher resolution and fine editing of HeyGen/Synthesia. On the technical side, it relies on Shengshu's self-developed TurboDiffusion inference acceleration framework and TurboServe streaming deployment engine, and superimposes attention optimization and model quantification such as SageAttention, SLA, SpargeAttention, etc., to reduce the delay of real-time video generation to the passing line on consumer-grade graphics cards. Some comments juxtapose it with MaineCoon Chat Mode of the same period, believing that the overall real-time video digital human product is still in the "pre-modern" stage, but the direction is clear.
-
The main controversy focuses on the management of experience expectations: there is a gap between the officially promoted "semantic perception, intention understanding, and corresponding comfort response" and the actual test. Voice control behavior is unstable on complex instructions, which can easily make users feel "fooled." The somatosensory shortcomings of half-duplex and higher latency in companion scenes have also triggered discussions on "whether virtual companionship is really ready." In addition, 540P resolution is still a hard limit for e-commerce live broadcasts and film and television promotions that require high-definition appearances; character details are not as controllable as 3D modeling solutions, which also means that it has temporarily given way to high-precision professional production.
-
It is suitable for: individual creators who want to be exclusive digital people at low cost, live broadcast teams who need to be online 24 hours a day, game developers who want to interact with NPCs in real time, companies who need brand digital people for customer service, and teams who want to quickly verify real-time companionship/education products. Not suitable for: professional film and television productions that pursue 4K film and television quality, projects that require complex multi-character scenes or 3D high-precision modeling. In terms of usage recommendations, individual creators should first use free credits to test the matching of actions and sounds before deciding whether to use it commercially; it is more stable for commercial teams such as e-commerce and education to directly use the enterprise API package. In terms of alternatives, low-latency products such as Doubao are available for pure voice companionship, and HeyGen and Synthesia are available for offline digital human short film production.
-
Vidu S1 transforms AI video from "shooting a short video" to "opening a real-time digital human call". Zero training on a single picture, voice-driven full-dimensional behavior, and unlimited time interaction do constitute a generational difference. However, delay, action controllability, and realism are still hurdles it must overcome before it can be accompanied by the public.
User Reviews
-
BillyBaker_2023—You can pinch people with just a picture, the threshold is really low. -
龙昊—The 540P image quality is indeed blurry. It’s okay to use it for video calls, but don’t even think about using it for serious filmmaking. -
钱怡凤—The simple command "raise hands" is smooth, the compound command is barely there, and complex movements are just confusing. -
郭丹丹—In experience, the delay is really not small. After you finish speaking, it will be stunned for a moment before speaking. It's fast to generate a video, but once it's used as a video call, this kind of pause is particularly annoying. -
傅彤晨—I recorded my own voice and wanted to customize the sound, but it didn't work at all. The male and female voices were randomly cut into the screen, which looked quite weird. I hope this can be fixed. -
Patrick.Gray—Tell it "I'm in a bad mood" and it will do soothing movements. In theory, sensing emotions is cool, but in reality there is still a gap between the ideal state promoted by the official, so don't expect it to be too smart. -
Virginia.Bell00702—After doing the math, I was a bit dissuaded: 3 points will be deducted for 2 seconds for real-time interaction, which is equivalent to about 2.8 yuan per minute, which is 160 yuan per hour. I spent more than 2,000 points and more than 1,000 seconds of connection to complete the entire multi-modal closed-loop evaluation. The free quota is enough for personal use, but if I really want to use it for e-commerce live broadcasts or companionship, I have to calculate the long-tail costs first, otherwise it will be burned out in a few minutes. It is recommended that the commercial team directly adopt the enterprise API package instead of using the free quota. -
AXmoo—An automated script was used to connect the API to run the stress test. The default role library was expanded from the earliest 6 to 11, and the iteration speed was very fast. After the two groups had a long conversation, the screen was basically fixed at 25 frames with zero jitter. Even after chatting for a long time, the faces did not fade. The engineering stability is indeed more stable than expected. I used Claude to get through the connection in more than an hour. I even invited Haaland and Wozniak to do stress tests in Chinese and English. To access it, you need to read the documentation and add the Alibaba Cloud ARTC SDK. The threshold is not low for pure beginners, but programmers who are willing to put in the effort can basically get started in a day. -
JRogersZ—It’s great that consumer-grade graphics cards can run 540P real-time generation. Small teams and individual creators can get started without stacking A100s, and the threshold for local deployment has been greatly reduced. -
ticklishgorilla525—Compared with HeyGen and D-ID, S1 is indeed a generational leader in depth of interaction and continuous dialogue, but with the 540P resolution, serious film and television-level production still has to rely on offline solutions, and the positioning must be clearly defined.