DeepMotion

用普通视频或文本提示生成 3D 角色动画的 AI 动作捕捉平台,无需动捕服或专业棚

In-depth Report

  • DeepMotion is a company that makes "AI video motion capture" available out of the browser. The core idea is to shoot a video with an ordinary camera and generate a 3D animation that can be directly imported into the game engine and digital human pipeline. The entire process does not require a motion capture suit, calibration camera or professional studio. Its two product lines—Animate 3D, which converts video to animation, and SayMotion, which generates animation from text—cover two creative methods, from material-driven to prompt word-driven. It is one of the most frequently mentioned “hardware-free motion capture” tools among independent developers, virtual anchors, and educational scenes. Its real strength lies in "lowering the threshold for professional motion capture to the floor." Its shortcomings are the depth ambiguity and jitter problems inherent in monocular vision, and complex movements still require manual frame editing.

  • DeepMotion Inc. was founded in 2014 and is headquartered in San Diego, USA. It is a company focusing on "AI motion intelligence". In the years since its establishment, professional motion capture was still dominated by inertial motion capture suits and optical marking systems such as Xsens. The investment for a set of equipment and sheds often cost hundreds of thousands of yuan, which was simply unaffordable for small and medium-sized teams. DeepMotion is betting on computer vision and deep learning: now that mobile phones and web cameras are popular enough, can the algorithm directly deduce 3D skeletal motion from a single 2D video? This route allowed it to stand in the first echelon when AI video motion capture broke out around 2020. The partners displayed on the official website include Intel, Nvidia, MIT, Samsung, Epic, Qualcomm, etc. The ecological niche is very accurate - it does not sell hardware, but embeds motion capture capabilities into other people's workflows in the form of SaaS and APIs. After years of operation, DeepMotion is already an independent company with ongoing revenue that relies on subscriptions and APIs to monetize, rather than a demo product that burns money through financing. This is not common among similar AI tools.

  • DeepMotion’s two major product lines have clear positioning. Animate 3D is a "video to animation": the user uploads a piece of MP4, MOV or AVI, and AI automatically identifies the key points of the human skeleton, back-calculates the joint angles, bone rotations and position changes, and then maps it to a standard 3D skeleton for export. It supports full body, hand, and facial tracking to be turned on at the same time (face and hand tracking each require additional points), and can track up to 8 people at the same time in a single video. Furthermore, it provides physical simulation, foot locking (Foot Locking), hand contact, motion smoothing (Motion Smoothing) and other options, as well as a so-called patented Rotoscope pose editor, which allows users to make frame-by-frame point corrections on the input video to specifically deal with tracking errors. SayMotion is a generative route for "text to 3D animation". The bottom layer is a MotionGPT class model based on prompt words. Users describe actions in natural language (such as "a sad man walking in the rain"), and the system generates multiple animation variations, and then uses tools such as Inpainting (generating local complementary frames), Merge (multi-segment splicing transition), and Refine (refining the original capture combined with prompt words) for secondary creation. In the official tutorial from the end of 2025 to 2026, the complete process of combining three segments of "Shaolin Kung Fu + Baseball Pitching + Basketball Dunk" into a multi-scene animation using video and text has been demonstrated, indicating that its workflow is moving from "single segment conversion" to "narrative-level arrangement". Both lines use cloud processing and browser operation. The exported formats cover FBX, BVH, GLB and MP4, and can be directly fed to Blender, Maya, Unity, and Unreal Engine. It also has built-in character creators for Ready Player Me and Avaturn, and supports uploading custom FBX/GLB/VRM models for bone redirection. Officials claim that the capture accuracy can reach up to 95%, and can reduce the hand K frame time of actions such as walking cycles by up to 90%.

  • DeepMotion uses dual-track billing of freemium plus animation credits, and Animate 3D and SayMotion are priced independently. The free version of Animate 3D provides about 60 seconds of animation quota per month and is limited to personal non-commercial use; the paid version ranges from Starter (about $15/month) to Freelancer/Pro (about $42–49/month, including commercial licensing and priority processing) to Studio (about $199/month, for teams and the highest priority tasks). The free version of SayMotion is 3 animation points per month, and the paid version is Starter at US$9/month, Professional at US$39/month, and Studio at US$83/month (all are annual prices). The points rule is 1 point = 1 second of animation or 1 3D pose. Turning on face or hand tracking adds 0.5 points per second for each segment. The API, Animate 3D API and Real-Time SDK for high-volume customers are all custom quotes and need to contact sales. It should be noted that there are discrepancies in the tier names and prices given by different sources (for example, some information lists Pro at US$48 and Studio at US$199.99, and some review sites list Starter at US$17 and Pro+ at US$149), which is speculated to be related to region, annual/monthly payment, and continuous price adjustments. The actual price is subject to the official website checkout page. Overall, its pricing is very friendly to light trials, but the points consumption and time limit for heavy use will become hidden costs.

  • Positive feedback is highly focused on "efficiency" and "cheapness". Jeasy Sehgal, a visual production professor at Georgia State University, said that compared to other software and hardware, DeepMotion is more friendly to independent creators, and facial capture is especially trouble-free; digital solutions expert Charles P. said that it has completely changed his game character animation pipeline; Corridor Digital artist Jordan A commented that its speed, accessibility and quality are "amazing", and it is a must-have tool for people who rely on motion capture work; Ruihan Zhang, a doctoral student at MIT Media Lab, used it in the AI ​​movie production hackathon. The most common scene mentioned by independent animators and XR creators is to use a video shot by themselves to make the characters pinched in VR move. Negative feedback focuses on the instability of output quality. Multiple independent actual measurements have mentioned that when the shooting conditions are not ideal, DeepMotion's output will show obvious "twitching, sliding, and feet not touching the ground": bones will be lost when limbs move quickly or when wearing loose clothes, tracking will collapse when a character is briefly blocked (such as being blocked while swinging), and plaid shirts will be misjudged in dance videos. Some users also complained about the fast consumption of points, too little free quota (some sources say the free quota is only 25 seconds, which is different from the official website's statement of 60 seconds), as well as the waiting and privacy concerns caused by cloud processing.

  • In the horizontal evaluation of AI video motion capture, DeepMotion is usually considered to be the one with the best balance of accessibility and animation friendliness, rather than the one with the highest accuracy. The multi-tool comparison in 2025 listed Move.ai as the choice with the highest accuracy for complex gestures, Quickmagic and Autodesk Flow Studio as budget choices that balance cost performance and quality, and Cascadeur and Marionette as offline solutions that are good at post-frame editing; DeepMotion and Plask were classified as "quick to get started, suitable for rough embryos and pre-previews, but have limited accuracy for complex animations". The original review pointed out that it "has noticeable twitching and sliding (obvious twitching and sliding)", the free version is extremely short. This type of evaluation points to an industry consensus: Monocular visual motion capture relies on computer vision to infer 3D from 2D, and there must be depth ambiguity - the result is correct from the camera perspective, but when placed in the three-dimensional space, the feet will float and the depth is not necessarily accurate. Therefore, the true positioning of DeepMotion is "preliminary rough capture + rapid prototyping". Key shots still rely on traditional marking systems or manual refinement. It forms a clear gradient division of labor with Move.ai (multi-camera high-fidelity), Rokoko Vision (free dual cameras), and Cascadeur (strong editing).

  • The biggest point of controversy is the “gap between precision propaganda and real experience.” Officially advertised up to 95% accuracy and 90% It is time-saving, but the conclusion reached by multiple independent tests when deliberately challenging boundary conditions (occlusion, loose clothing, fast movements) is that "it is relatively stable under ideal conditions, but shakes significantly, feet do not touch the ground, and usability is limited by many shooting requirements." It is highly bound to the best practices listed in the user manual (the camera is fixed and parallel to the actor, the whole body is in the frame, the character and the background are high, knees and elbows are not covered with loose clothing, and there is no occlusion) - essentially "garbage in, garbage out." Secondly, there are two risks brought by dependence on the cloud: First, privacy. Uploading character videos to third-party servers is a compliance risk for sensitive projects; second, workflow bottlenecks, cloud computing power will become a stuck point when processing large libraries of videos in batches. Coupled with the points system's restrictions on heavy use, face/hand tracking is still in the enhancement stage that requires additional payment, and the upper limit and stability of multi-person tracking are still being polished. These are the boundaries of "it can be used, but don't expect to get it right at once".

  • DeepMotion is most suitable for three types of people: independent game developers and small teams (you can get high-quality animation assets without investing in equipment), virtual anchors and short video creators (live performances drive avatars), and education and research institutions (use text + video to quickly produce teaching animation materials). It’s also suitable for rapid prototyping of film and TV previs, VR/AR and Metaverse content. Scenarios that are not suitable: key shots that require extremely high movement precision and cannot accept manual frame editing; projects that are highly sensitive to data privacy; and processes that require complete offline and local control (such requirements should be looked at Cascadeur or Marionette). If the budget allows and you are pursuing complex pose fidelity, you can evaluate Move.ai in parallel; if you are just doing early blocking, the free files of Plask and Rokoko Vision are also worth a try.

  • DeepMotion has liberated "professional motion capture" from hardware and studios to browsers, and is one of the most solid players in the popularization of device-free 3D animation. It is very cool to use it as "pre-production capture + rapid prototyping", but don't underestimate its monocular accuracy. Leave a budget for manual frame editing for key actions.

User Reviews

  • 头像
    Danielle30
    给大学动画课当教学工具很合适,学生不用买动捕服也能体验从视频到 3D 动画的完整流程,作业产出明显比纯理论课丰富。就是免费额度对全班同时用有点紧张,我们最后买了个教室账号才兜住。

  • 头像
    Denise_RodriguezIII
    卡成 PPT 的时候是真的急。

  • 头像
    gr3hnerdb
    跟 Move.ai 比精度确实差点意思,复杂姿态和多人互动 DeepMotion 还嫩,但胜在便宜、浏览器直接跑、不用多相机布置。我们评估一圈最后选它,就是因为团队就两个人,上多相机方案成本划不来。

  • 头像
    浮生_21
    SayMotion 写提示词生成动画用来做 early blocking 特别顺,跟 Animate 3D 的视频捕捉正好互补,前期创意验证效率很高。缺点也有,生成的动作有时候偏「标准」,情绪和细节还得自己加,另外中文提示词理解一般,用英文更稳。

  • 头像
    EL_ngu
    积分每月清零不累积,重度用户得算好账。

  • 头像
    Ronald_SandersII85
    把 DeepMotion 接到 Unreal 重定向到 MetaHuman 的那套流程跑通后,做短视频数字人效率直接起飞,一个下午能出好几条带身体+面部+手指的动画。前期踩坑主要在骨骼匹配和导入设置,官方教程跟一遍就顺了,比找外包划算太多。

  • 头像
    Carol_Barnes520255
    做独立游戏最缺的就是动画资源,以前要么买 Mixamo 要么硬手 K,现在手机拍段视频传上去就能拿到 FBX,导入 Unity 直接能用。走路循环、挥手这类重复活儿基本被它承包了,省下的时间够我多调两关,对 solo 开发者来说这工具简直是外挂。

  • 头像
    SWalkerX
    回不去了,手 K 帧太痛苦。

  • 头像
    Nathan_Martin168
    太香了,免设备直接出动画。

  • 头像
    Dylan_Bell_99
    整体 4 星,扣的那颗星给积分制和偶尔抖动。它确实不是万能的,复杂动作和多人互动还得人工精修,但作为前期捕捉和快速原型工具,性价比在同类里数一数二,新手完全可以先从免费档摸起,够用再升级。

  • 头像
    8r6jhlx_gt7
    云端处理高峰期要排队,长视频等得有点久,急活儿建议错峰或者上付费优先队列。有次赶版本交一段两分钟的动画,免费档卡了快一小时才出,后来咬牙升了 Pro 档走优先处理,速度立刻上来,这笔钱在 deadline 前还是值得的。

  • 头像
    月光135
    做虚拟主播的可以冲,摄像头实时驱动 3D 形象,延迟低、还带面部和手势,直播间效果比纯 2D 立绘生动太多。

  • 头像
    Frances.Hill_2021
    标牌说最高 95% 精度,实测理想条件下还行,一到遮挡、宽松衣服、快速动作就露怯,别被宣传数字带偏。

  • 头像
    JoshuaMyers_77
    视频拍得一般的时候手脚会抖、脚不沾地,得开 Rotoscope 逐帧修,复杂动作还是得人工兜底。

  • 头像
    RCarterSr
    我们工作室拿它做 NPC 和群演的预演动画,量大管饱还便宜,关键镜头才上传统动捕,性价比拉满。

  • 头像
    郑悦杰
    面部和手部追踪要额外收费,0.5 积分/秒,用多了肉疼。

  • 头像
    AnaGuevara
    免费档一个月才 60 秒,试两下就没了,有点抠。

  • 头像
    NoahJenkins
    想要高质量还是得自己拍好素材,相机固定、全身入镜、紧身衣、背景反差大,垃圾进垃圾出一点不假。

  • 头像
    blueostrich338
    导出格式挺全,FBX/BVH/GLB/MP4 都有,Blender、UE、Unity 通吃,管线接入基本没阻力。我们项目是 Unity 加自研角色,上传 FBX/GLB 后重定向到自己的骨骼命名规范很顺,连 Mixamo 那种要先绑骨骼的步骤都省了,对中小团队真的很友好。

  • 头像
    琉璃_15
    用 Animate 3D 把一段功夫视频转成游戏角色动画,前后不到十分钟,比手 K 帧快太多了。以前这种打斗动作光调骨骼就得好几个小时,现在上传、选角色、等几十秒就能拿到可导入的动画,虽然脚部偶尔要修,但粗胚阶段省下的时间非常可观。

  • 头像
    枫叶129
    用 Rotoscope 姿态编辑器描点修正能赚回积分,官方这招挺聪明,既帮用户改动作又顺手喂数据。

  • 头像
    崔婷玉
    DeepMotion 真把动捕门槛打到地板价了,独立小团队福音。