Taichu

中国科学院自动化研究所研发的full-stack 国产化multimodal large model

In-depth Report

  • Zidong Taichu is a large domestic multi-modal model jointly developed by the Institute of Automation of the Chinese Academy of Sciences and the Wuhan Institute of Artificial Intelligence. As the world's first 100-billion-parameter tri-modal pre-training model, Zidong Taichu realizes cross-modal understanding and generation of images, texts and sounds. The latest version 4.0 achieves a major breakthrough from "passive analysis" to "active thinking", supporting innovative functions such as 180-minute long video understanding and multi-modal in-depth reasoning. It has in-depth applications in fields such as smart manufacturing, smart medical care, and embodied intelligence.

  • Zidong Taichu was developed by the Institute of Automation of the Chinese Academy of Sciences and released version 1.0 in July 2021. It is the world's first three-modal pre-training model with 100 billion parameters of image, text, and audio. Version 2.0 will be released in June 2023, adding support for video, 3D point cloud and other modalities. Version 3.0 will be released in November 2024 to enhance hybrid understanding capabilities. Version 4.0 will be released in September 2025, achieving a leap forward in "active thinking" and introducing a human-like cross-attention mechanism. The platform is launched based on the rich knowledge accumulation and experience of the "Zidong Taichu" large model of the Institute of Automation, and is oriented to industry AI applications.

  • Zidong Taichu provides rich multi-modal AI capabilities. In terms of multi-modal deep reasoning, it supports the cognitive ability of seeing, learning and thinking while realizing the leap from passive analysis to active thinking. The ability to think with pictures supports fine-grained operations such as translation, amplification, rotation, positioning, enhancement and reconstruction of images. Complex reasoning ability can handle professional problems such as mathematical reasoning and calculation. Long video understanding enables for the first time in-depth understanding of 180-minute long videos and accurate answers within seconds. In terms of full-modal support, cross-modal content generation such as "pictures to produce sounds", "pictures to produce text" and "sounds to produce pictures" in three modes of picture, text and sound is realized, and more modal data such as 3D point clouds, videos and signals are integrated. Supports document analysis (maximum 10MB), multiple rounds of question and answer dialogue, audio authenticity identification and other functions.

  • Zidong Taichu's technical features include: unified semantic representation technology, which maps multi-modal data such as images, text, voice, video, 3D point clouds, signals, etc. into the same semantic space; version 4.0 introduces a human-like cross-attention mechanism to realize active thinking; all-modal open access realizes full-modal access to structured and unstructured data; grouped cognitive encoding and decoding enables full understanding and flexible generation of multiple data information; cognitive-enhanced multi-modal correlation technology effectively integrates cognitive-enhanced multi-modal correlation of multi-tasks. In terms of training efficiency, 100% supervised learning effect can be achieved based on 5%-10% of data annotation. It supports multi-task joint learning under unsupervised conditions and rapid data migration in different fields. The model supports lightweight deployment and inference acceleration, and is suitable for different hardware environments.

  • Zidong Taichu has been applied in many industries. In the field of intelligent manufacturing, the intelligent welding precision cooperated with Huagong Technology reaches 0.02 mm, surpassing that of a ten-year-old master. It only takes 43 seconds to weld the entire vehicle. It captures the weld gap, misalignment, etc. in real time, generates the optimal path in milliseconds, and supports 25 types of intelligent welding processes. In the field of smart medical care, it helps Jiuzhoutong manage tens of thousands of medical devices and consumables. The inventory inventory time is reduced from 3 days to 4 hours, and the efficiency is increased by 30 times. In terms of surgical assistance, it is deployed on the neurosurgery robot MicroNeuro, which integrates multi-modal information such as vision and touch in real time during surgery. In the field of embodied intelligence and low-altitude economy, five robot vocational skills training schools have been built in Wuhan, Foshan, Qingdao and other places to empower intelligent decision-making and path planning of low-altitude aircraft such as drones.

  • As a representative of domestic multi-modal large models, Zidong Taichu has filled the gap in the domestic general multi-modal field. Compared with foreign models, Zidong Taichu’s core advantages lie in full-stack localization and depth of industrial application. The platform has formed mature application cases in vertical fields such as smart manufacturing and smart medical care, proving the practical value of the model. The “active thinking” ability of version 4.0 is its differentiated competitive advantage.

  • Zidong Taichu is suitable for the following user groups: enterprise users, industry customers who need multi-modal AI capabilities; developers, technical personnel who develop applications based on large models; researchers, scholars engaged in multi-modal AI research; ordinary users, individual users who experience multi-modal interaction.

  • As a full-stack domestic multi-modal large model, Zidong Taichu has innovative advantages in multi-modal in-depth reasoning, active thinking, and 180-minute long video understanding. It has in-depth accumulation in intelligent manufacturing, smart medical and other industrial fields, and is an important representative of the development of domestic large-scale models. For users and enterprises that require multi-modal AI capabilities, Zidong Taichu is an option worth paying attention to.

User Reviews

  • 头像
    Aaron_Kim_Max
    紫东太初的长视频理解能力很强,180分钟视频也能精准回答问题。

  • 头像
    PaulRoberts
    作为国产大模型,能做到这个水平已经很不容易了,继续加油!

  • 头像
    JeffreyWerner
    在智能制造领域的应用案例很惊艳,焊接精度0.02毫米超越老师傅。

  • 头像
    Sam_anthaWood
    多模态能力很全面,图文音视频都能处理。

  • 头像
    cegevmxp
    4.0版本的主动思考能力是一大亮点,和国外模型有一战之力。

  • 头像
    Laura_Gonzales2
    医疗领域的库存管理效率提升30倍是真的强!

  • 头像
    TuckerKin_g
    全栈国产化很重要,信息安全问题不用担心的。

  • 头像
    Shirley.Perez330
    智慧医疗应用案例很实际,不是概念产品。

  • 头像
    SamuelKujala
    具身智能的机器人培训很有前景,支持!

  • 头像
    CArey
    期待更多行业应用落地,继续优化模型。

  • 头像
    mrqsc0
    和华为昇腾适配进展怎么样了?

  • 头像
    de5pgb
    有API可以申请试用吗?

  • 头像
    Jacqueline47
    开源程度怎么样?个人开发者能使用吗?

  • 头像
    Denise.Stephens_2020580
    和GPT-4V对比哪个更强?

  • 头像
    海浪_27
    中科院出品必属精品!

  • 头像
    DiamondHands493
    3D点云理解的准确度怎么样?

  • 头像
    Sophia.Williams_2020
    轻量化部署支不支持消费级显卡?

  • 头像
    MargaretMendozaII
    语音生成能力有待提升。

  • 头像
    潘博星
    文档解析最大支持10MB很实用。

  • 头像
    夏风936
    期待5.0版本的表现!