LPM 1.0
17 billion parameter real-time full-duplex conversation AI model for video character performance
In-depth Report
-
LPM 1.0 (Large Performance Model) is a large-scale video character performance model released in April 2026 by Anuttacon, an AI company founded by MiHoYo founder Cai Haoyu. The model uses a 17 billion parameter diffusion Transformer architecture and is designed for real-time generation of character videos that can speak, listen, react, and maintain consistent identities over long periods of interaction. LPM 1.0 is the first model to simultaneously achieve true full-duplex conversational video generation, with a latency of only 0.35 seconds (one-third of competing products), and can maintain zero identity drift for more than 45 minutes.
-
LPM 1.0 was developed by the Anuttacon team, an AI team composed of more than 20 researchers and closely related to MiHoYo founder Cai Haoyu. The team released LPM 1.0 in April 2026, and published a technical paper on arXiv during the same period (No.: arXiv:2604.07823). LPM 1.0 is positioned as a large-scale performance model for video character performances, focusing on solving the visual generation problems of dialogue AI agents, virtual anchors and game NPCs. This model represents a new breakthrough in the field of video generation, pushing traditional one-way video generation into a new paradigm of two-way full-duplex dialogue. From the perspective of technology inheritance, LPM 1.0 benefited from the early accumulation of the ByteDance research team and achieved a qualitative leap on this basis. The Anuttacon team noted in its public technical report that the model uses a 43-page technical document to describe its 15 technical innovations in detail.
-
LPM 1.0 implements several industry-leading breakthrough features. In terms of real-time performance, LPM 1.0 achieves an end-to-end delay of only 0.35 seconds, which is only one-third of competing products. This means users can achieve a nearly seamless conversation experience with no apparent wait. In terms of full-duplex dialogue, LPM 1.0 is the first full-duplex dialogue video generation system that can achieve both "speaking" and "obedient" states at the same time. Unlike traditional models that can only generate character speech, LPM 1.0 can generate reactive listening behaviors while the character speaks, including eye contact, expression changes, nodding responses and other micro-performances. In terms of identity consistency, LPM 1.0 adopts multi-granularity identity conditionalization technology, which can maintain zero identity drift during long interactions of more than 45 minutes. This is recognized as one of the most difficult problems to solve in the industry, and LPM 1.0 provides a satisfactory solution for the first time. In terms of multi-modal control, LPM 1.0 uniformly supports three control signals: text, audio and image. Users can control the character's performance behavior through any combination. In terms of zero-shot generalization, LPM 1.0 supports characters of any style without the need for fine-tuning for specific styles. Including realistic styles, animation styles, 3D rendering styles and non-humanoid characters can be naturally generated. In terms of output specifications, LPM 1.0 supports two resolutions, 480P and 720P, with a frame rate of 24fps for real-time streaming output.
-
As of now, the specific business model of LPM 1.0 has not been made public. Judging from the fact that the model has just been released and the technical report has been made public, the model may currently be in the early research or testing stage. Referring to industry practices, possible future business models of LPM 1.0 include: enterprise-level services that charge based on the number of calls through the API interface; SDK authorization for developers; customized solutions for the gaming or live broadcast industry; and personal subscriptions for virtual anchors and content creators. As part of the MiHoYo ecosystem, LPM 1.0 may first be used in MiHoYo's own game products in the future to provide real-time dialogue capabilities for its virtual characters.
-
Since LPM 1.0 was just released in April 2026, the number of public user comments is limited, but judging from feedback from the technical community, the industry speaks highly of it. In the technology community, LPM 1.0 is called the "Game Changer" in the field of video generation. Researchers are particularly concerned about its achievement of zero identity drift within 45 minutes, believing that this solves an industry pain point. The implementation of full-duplex dialogue has received strong attention from the virtual anchor community. Compared with traditional methods that require a lot of post-production to achieve interactive effects, LPM 1.0 can generate the reactions of both parties to the conversation in real time, significantly lowering the content production threshold for virtual anchors. Game developers have expressed concerns about LPM 1.0's ability to generate NPC dialogue in real time. If it can be applied to game NPCs, it will bring revolutionary changes to open world games.
-
The release of LPM 1.0 has attracted widespread attention in the AI video generation industry. From a technical perspective, the 17B parameter diffusion Transformer architecture adopted by LPM 1.0 is one of the largest video character performance models in the industry. Compared with competing products such as LiveAvatar, OmniHuman, and Kling-Avatar-2, LPM 1.0 has obvious advantages in real-time performance and identity maintenance. From an industry impact perspective, LPM 1.0 represents an important paradigm shift in video generation from one-way output to two-way interaction. This transformation means that AI can not only "speak" but also "listen", which is the basic element of natural dialogue between people. The Bytedance research team played an important role in the development of LPM 1.0, which shows the continued investment of large technology companies in the field of video AI.
-
So far, there has been no major controversy in the LPM 1.0 project itself. However, as a cutting-edge technology product, users need to pay attention to the following risk factors. In terms of technology maturity risks, LPM 1.0 has just been released, and it will take time to verify the actual deployment effect and large-scale application stability. Whether its claimed performance indicators can be achieved in actual scenarios remains to be seen. In terms of commercialization risks, the business model and pricing strategy of the model have not yet been clarified, and there may be risks of higher usage costs in the future. In terms of legal and ethical risks, real-time full-duplex conversation video generation technology may be abused for improper purposes such as deep forgery, and attention needs to be paid to the development of relevant regulatory policies.
-
LPM 1.0 is suitable for the following user groups: virtual anchors can achieve natural real-time interaction with audiences through LPM 1.0, greatly reducing the content production threshold for interactive live broadcasts; game developers can use LPM 1.0 to empower NPCs with real-time dialogue capabilities, creating a more immersive gaming experience; AI dialogue product developers can add video output capabilities to dialogue AI agents to achieve more natural human-computer interaction; content creators can use LPM 1.0 to quickly generate high-quality dialogue video content. For ordinary users, it is recommended to wait until the model is commercialized and a more user-friendly interface is provided before trying it. For technical researchers, you can learn more about its technical implementation through arXiv papers.
-
LPM 1.0 is a milestone product in the field of video character performance. Its realization of full-duplex dialogue and 45 minutes of zero identity drift represents a major breakthrough in the industry. Although the path to commercialization is not yet clear, its technical potential deserves attention.
User Reviews
-
云朵_2—Full-duplex conversations are the real breakthrough. In the past, other models could only generate speech, and the reaction had to be done later, but now it can finally be generated in real time. -
BruceBerg—45 minutes of zero identity drift is outrageous! The models used before started to fall apart within a few minutes. LPM 1.0 directly solved the problem that has plagued the industry for a long time. -
Rachel.KingSr—What is the concept of 0.35 second delay? The delay is almost imperceptible! In the past, using other delays would give me a sudden death. -
Betty.Hughes_7—The 17 billion parameters are indeed not in vain, and the effect is better than all models used before. -
马贞睿—Realistic animation 3D can be generated without separate fine-tuning. This is so cool. -
ShirleyWood_2024—I'm not surprised that MiHoYo did this. Their virtual character technology has been accumulated for a long time. -
LItic—If game NPCs can use this, the open world experience will be directly upgraded to a higher level. -
Samantha_Simmons_77—Is ByteDance also involved? It seems like a strong alliance. -
Jason_Phillips_Max—I am afraid that the threshold for virtual anchors will be lowered by this model in the future. -
枫叶_8—Cai Haoyu personally led the team, and what came out was indeed different. -
3f71zhr—The first model to achieve full-duplex conversation at the same time, this title is quite domineering. -
竹影344—LiveAvatar and OmniHuman were crushed. This wave is really a slap in the face. -
DMartinez_2023—The name Zhuhuo Technology is very interesting, and I hope it will become popular in the field of AI. -
BrandonHernandezQ—There are 15 pictures on page 43 of the technical report. Are you serious about investing in this? -
JudithRoberts_88—Multi-modal control is also unified and should be very convenient to use. -
PDiaz_7715—As soon as this model came out, it felt like the industry was going to change. -
JackSanchez_X14—When can I experience it? Waiting. -
SWilliams520771—48OP/720P 24fps real-time output, the picture quality is also top-notch. -
梅花_2—I have read the paper on arXiv, and there is indeed something in the architecture.