Google Gemini 4 Ultra
Google DeepMind 于 2026 年 5 月发布的第四代旗舰多模态大模型,支持千万级 Token 上下文、原生多模态生成与 Agentic 自主工作流,搭载第七代 Ironwood TPU
In-depth Report
-
Google Gemini 4 Ultra is the fourth-generation flagship multi-modal large model released by Google DeepMind at the Google I/O conference on May 19, 2026. As the pinnacle of the Gemini series, Gemini 4 Ultra adopts the "Agentic-first" design concept and achieves a comprehensive leap in terms of context length, multi-modal generation capabilities, autonomous reasoning, and cross-application task execution. It is regarded as the core engine of Google's move from "AI assistant" to "AI operating system".
-
The biggest breakthrough of Gemini 4 Ultra comes from its context window - it supports up to 10 million Tokens (the actual subscription version provides 2 million Tokens), which is 10 times that of the previous generation Gemini 2.5 Pro and one of the largest context windows in the industry at the time. This means it can handle the entire text of an entire novel, hours of video/audio transcription, or architectural analysis of an entire code repository in a single inference. For researchers, hundreds of papers can be directly fed into the model at once for comprehensive analysis and cross-reference; for developers, the business logic and call links of the entire project can be understood in one session.
-
In terms of benchmark testing, Gemini 4 Ultra achieved a score of approximately 84.6% on ARC-AGI2 (Abstract Reasoning Challenge), significantly surpassing GPT-5.5 (approximately 7x%) and Gemini 2.5 Pro (approximately 52%), demonstrating a qualitative improvement in reasoning capabilities under new constraints. The MMLU score was 92.5, and the long text retrieval accuracy (Needle In A Haystack) was as high as 95.2%, both ranking in the first echelon at the time.
-
Gemini 4 Ultra adopts a sparse MoE (Mixture of Experts) architecture and introduces an innovation called "knowledge activation gate" by Google - the model will first determine whether the user's problem requires calling an external or local knowledge base, and then decide how many MoE experts to activate. According to Google official data, compared with 2.0 Ultra, the response speed when calling external knowledge has increased by 47% and the accuracy has increased by 12% under the same computing power.
-
This model is a native multi-modal design that not only supports the input and output of text, images, and audio, but also supports video generation (based on the Veo 3 model), music generation, and spatial perception output (adapted to AR/robots) at high fidelity for the first time. Gemini 4 Ultra's video generation capabilities can generate high-definition coherent short films in 30 seconds, with built-in storyboarding capabilities that allow users to control the creative process frame by frame.
-
Gemini 4 Ultra runs on Google's seventh-generation TPU - Ironwood (TPU v7). Ironwood's peak performance is approximately 10 times that of TPU v5p, and its single-chip inference/training performance is 4 times higher than that of Trillium (TPU v6). This deep collaboration at the hardware level allows Gemini 4 Ultra to maintain controllable inference latency even in extreme scenarios where tens of millions of Token contexts are processed.
-
From the product matrix, the Gemini 4 series is divided into four levels: Gemini 4 Nano (end-side core, memory usage after INT4 quantization is less than 600MB, latency as low as 35ms, adapted for offline operation of Android 17 devices); Gemini 4 Flash (edge server design, low latency optimization); Gemini 4 Pro (balanced, supports 128K Token context, oriented to general scenarios); and Gemini 4 Ultra (flagship level, for extreme performance needs).
-
The application scenarios driven by Gemini 4 Ultra cover Google's entire product ecosystem. In the Gemini app, it serves as the default model; in Google Search's AI Mode, it provides multi-step reasoning and conversational answers with source references; in Android 17's Gemini Intelligence system-level AI layer, it drives autonomous workflows across applications (a single command such as "find my course syllabus from Gmail, identify textbooks, add to cart" can complete multi-step operations); and on Android XR smart glasses, it provides real-time translation, visual understanding and navigation capabilities. Gemini 4 Ultra also introduces the “Gemini Spark” intelligent agent platform. Spark is a cross-application fully automated task agent that can obtain information from Gmail, operate shopping carts, complete orders and other operations without manual instructions from users. This is a key step for Google to shift AI from "passive response" to "active execution."
-
In terms of subscription plans, Google provides tiered pricing: the free version is limited to standard quota and context 32K Token; Google One AI Premium ($19.99/month) provides 1M context, 4x quota, Workspace integration such as Gmail/Drive/Docs, and 2TB storage; Google AI Ultra ($99.99/month) provides 5x quota, Deep Think inference mode, Gemini Spark agent, and YouTube Premium rights and interests. In addition, there is a more advanced Gemini Ultra subscription for professional users ($249/month, 2M context, Veo 3 unlimited video generation, 30TB storage). Gemini 4 Ultra brings significant improvements in AI capabilities in multiple dimensions, but there are also limitations worth paying attention to. The first is pricing—the highest-end subscription, $249/month, has a higher threshold for individual users and is mainly targeted at enterprises, research institutions, and heavy creators. Secondly, in pure code development scenarios, some reviews point out that Claude Max and GPT-5.5 still have advantages on a specific coding benchmark (SWE-Bench). In addition, the actual application scenarios of tens of millions of Token contexts are still relatively limited, and 1M-2M Tokens are more than enough for most users. The release of Gemini 4 Ultra also triggered a new round of discussions about the boundaries of AI capabilities and security governance. Google simultaneously announced the Safety Core framework to deal with potential risks brought about by the jump in model capabilities.
User Reviews
-
逍遥_10—刚升级了 Gemini 4 Ultra,200万上下文真的太夸张了,直接把整个项目代码库扔进去,它居然能跨文件分析依赖关系。 -
SatoshiSnakeJones—Deep Think 模式确实强,但一次对话直接烧掉我 40% 的配额,肉疼。日常用 Pro 就够了,Ultra 适合真要啃硬骨头的时候。 -
Cynthia.Morgan_7—用了 Ultra 两周,最深的感觉就是:Google 生态整合是真香。在 Gmail 里直接让它总结邮件、在 Docs 里改文档、在 Sheets 里写公式,完全不用切窗口。如果你本来就生活在 Google 全家桶里,这个订阅性价比确实高。 -
IRodriguez_20244—我专门试了一下它处理超大文档的能力,把一本 400 页的电子书扔进去让它做章节摘要,12 秒出结果,每个章节的关键论点都抓得很准。以前用其他模型分章节处理至少得来回切十几次上下文,现在一次搞定,效率提升不是一点点。 -
DennisMartin_Plus—吐个槽:Gemini Ultra 的 Deep Research 跑出来的报告有时会漏引用来源,核查起来反而更费时间。功能是好功能,但「可验证性」还得加强。 -
Kathryn.Wilson_2022—Gemini 4 Ultra 的视频生成比想象中好,30 秒生成了一个产品宣传短片,连贯性不错,但细节经不起放大看。商用还得再迭代几版。 -
Cynthia_RichardsonSr—把 Gemini 4 Ultra 接入了我司的客服系统,上下文理解能力比之前用的模型强很多,能记住用户前面提过的需求,不用每次都重新说。 -
TOrtizK—Ironwood TPU 加持下的推理速度确实快,处理 100 页 PDF 基本是秒级响应。但说实话,大部分场景我用 32K 上下文的免费版也够了。 -
CRuizSr—我是做科研的,经常需要同时读十几篇论文做交叉分析。Ultra 的 200 万上下文对我来说是刚需,以前要花一整天的事现在一小时搞定。 -
BeverlyThompson_Plus—Gemini Spark 智能体试了一下,让它「从 Gmail 找到上周的报价单,提取价格,做成表格发给我」,居然真的跑通了。虽然中间卡了一次,但方向是对的,期待后续优化。 -
Douglas_Coleman_2020—纯代码场景我觉得 Claude 还是强一点,Ultra 写 Python 还行,但写 TypeScript 的时候偶尔会脑补不存在的 API。不过如果是做 Google Cloud 相关的开发,Ultra 对 GCP 生态的理解明显更深,算是有舍有得吧。 -
Christina_Simmons168—SEA 地区用户表示,最近 Ultra 的响应延迟明显变高了,不知道是海底电缆问题还是服务器负载问题。凌晨用倒挺快。 -
SLewis—知识激活门这个设计挺聪明的,简单问题不调大模型,省算力又快。DeepMind 在工程优化上确实有一套。 -
RClarkQ1—把一份 200 页的招股书扔进去,让它找出所有风险披露条款并对比往年变化,五分钟出结果。这要是人工做至少一整天。$20 月费光这一单就回本了。 -
Jerry_BellIII—Android 17 上的 Gemini Intelligence 确实好用,长按电源键直接说「帮我订下周五去上海的火车票」,它自己打开 12306、选车次、填信息,就差最后一步付款了。这种跨应用的 Agent 能力,确实让手机从一个被动工具变成了主动助手。 -
IRISnetProHall007—今天用 Gemini 4 Ultra 做了一个竞品分析报告,从网上抓了 5 家竞品的公开信息,加上我自己的产品文档,它自动生成了一个带对比表格的完整报告。最惊艳的是它还能自己配图,Imagen Ultra 4 生成的对比图直接可以用在 PPT 里。虽然有些数据需要人工核实,但整体质量已经能省掉 80% 的案头工作了。 -
onnxcu—Gemini Code Executor 2.0 支持跨语言协作编辑挺惊艳的,我一个 Python 脚本写到一半让它在里面跑 SQL 查询,结果直接出表。 -
5vjig78g1—Imagen Ultra 4 生成的图片真的比 Midjourney 还真实,尤其是人像,皮肤质感和光影处理太自然了。商用素材可以直接用。 -
LawrenceVasquez_66—个人感觉 Ultra 的中文能力比 GPT-5.5 好,写中文文案的时候用词自然很多,不会那种翻译腔。 -
MsEeviTuomala_x—试了一下 Gemini 4 Ultra 的音频理解能力,把一段 2 小时的会议录音扔进去,它自动生成了会议纪要和待办事项,发言人都能区分,牛。 -
Zachary_Allen—ARC-AGI2 84.6% 确实厉害,但日常使用中感知不强。绝大多数用户的问题根本到不了需要这种抽象推理的级别。 -
珊瑚611—升级到 Ultra 之后最大的变化其实是配额焦虑减轻了,不用再精打细算每次对话用了多少额度。5x 额度够我随便造。