Gemini 3.6 Flash
Google DeepMind 面向智能体时代的高效主力模型,在编程、知识工作与多模态任务上更强、更省 token
In-depth Report
-
Gemini 3.6 Flash is the main model officially released by Google DeepMind on July 21, 2026. It is positioned as an "efficient workhorse" for the era of intelligent agents. It has improved on programming, knowledge work and multi-modal tasks compared with the previous generation 3.5 Flash. At the same time, it has reduced the output token usage by about 17% and the price has also been slightly reduced. However, this generation release did not bring the long-awaited flagship Gemini 3.5 Pro. In addition, the feedback from the first batch of developers was polarized, and the community's reputation dropped significantly. For enterprises and developers who need to use AI at a large scale and at low cost, it is still a cost-effective choice, but if you are looking for cutting-edge reasoning capabilities, you may be disappointed.
-
Gemini 3.6 Flash was developed by Google DeepMind and is a member of the Gemini 3 series. The technical route is directly based on the previous Gemini 3.5 Flash. Google chose to launch three lightweight models this time - 3.6 Flash, 3.5 Flash-Lite and 3.5 Flash Cyber for security scenarios - instead of the flagship Pro, which in itself sends a clear signal: the current focus is on providing customers who build AI agents on a large scale with more efficient, lower latency, and more reliable underlying capabilities. It is worth noting that the real flagship Gemini 3.5 Pro has been delayed multiple times. According to Bloomberg in mid-July, the 3.5 Pro has been delayed due to difficulty meeting internal performance indicators and is currently only open to partners for testing. Logan Kilpatrick, Google's DeepMind product lead, said that the team is testing 3.5 Pro with partners and hopes to "launch it as soon as possible" and has launched the largest Gemini 4 pre-training program in the company's history. Against the background of accelerated iterations by competitors such as OpenAI and Anthropic, Google has temporarily focused on the Flash line, which many observers interpreted as "using cheap Flash to stabilize the market."
-
Gemini 3.6 Flash is a native multi-modal inference model that supports text, images, audio, video and PDF as input and output as text. It has an input context window of up to 1 million tokens and an output upper limit of 64,000 tokens. It enables medium-intensity thinking capabilities by default. The built-in tools are quite complete: function calling, structured output, code execution, computer use (preview), search grounding, file retrieval, URL context, and Google map grounding, which can be used for complex agent workflows out of the box. Compared with 3.5 Flash, 3.6 Flash focuses on "talking less nonsense": by optimizing the thinking chain, the number of reasoning steps, dialogue rounds and tool calls required for multi-step workflows are reduced, alleviating self-entanglement in the execution cycle. The quality of code generation is improved, reducing unnecessary changes and debug loops; instructions are followed better, and files are less mistakenly modified during diagnostic tasks; multi-modal and spatial reasoning are enhanced, and performance is more stable in chart interpretation, visual blueprint conversion, and multi-element web layout generation. It also tends to run diagnostic scripts before making modifications, improving the accuracy of complex tasks at the expense of more exploratory steps on simple front-end tasks. However, the visual performance of the front-end code is controversial: Officials admit that human evaluators prefer the visual layout and style of the previous generation, and recommend that users proactively provide clear design guidelines to compensate.
-
Gemini 3.6 Flash uses standard API pay-as-you-go billing: input $1.50/million tokens, output $7.50/million tokens (3.5 Flash is 1.50/9.00, output cost is reduced by about 17%); input with cache is as low as $0.15, and cache storage is billed at $1.00/million tokens/hour. There is also a free file (limited usage). The 3.5 Flash-Lite released in the same batch is cheaper, with an input cost of $0.30, an output of $2.50, and a generation speed of 350 tokens per second. It focuses on high throughput and low latency tasks, and supports switching inference gears. 3.5 Flash Cyber is a specially fine-tuned security model used to discover and repair network security vulnerabilities, and is only open to government agencies and trusted partners in limited quantities. For companies that use a large number of AI, the token cost reduction of 3.6 Flash is often more realistic than the improvement of a few percentage points on the benchmark.
-
Positive feedback focuses on "fast" and "saving". In the actual test, 3.6 Flash was about one-third faster than 3.5 Flash on logic and calculation questions, and the token consumption was significantly reduced: the same seat arrangement constraint question, 2338 tokens vs. 3707 tokens; the AA apportionment question was completed in 9.4 seconds, while the 3.5 Flash took 12.4 seconds. In terms of knowledge work such as document parsing, chart data analysis, and report drafting, customers such as Hebbia and Harvey have received outstanding feedback, and Figma and JetBrains have also publicly endorsed it. Use it to make interactive web page prototypes, and the completion level is enough for people who do not understand the front end. Negative feedback is equally harsh. The first batch of developers found that the front-end code generation capabilities have regressed: the component nesting is confusing, the interaction logic is broken, and the code cannot be run directly. Some developers left lamenting that "this is the worst result I have ever seen." Spatial reasoning ability has declined, and image texture generation is occasionally acceptable but rendering bugs often occur. Netizens also complained about the outrageous choice of Chinese words, confusing memory, calling the wrong tools, poor image and video quality, and incompatible instructions. On the intelligence index of the third-party Artificial Analysis, 3.6 Flash scores the same as 3.5 Flash (about 50 points), lagging behind Claude Fable 5, GPT-5.6 Sol, Kimi K3, Grok 4.5, GLM-5.2 and other opponents. Some users bluntly said that it "saves tokens, but does not preserve IQ."
-
The picture given by third-party benchmarks is "partly leading, but overall mediocre." On agent/code tasks such as SWE-Bench Pro (58.7%), DeepSWE v1.1 (49%), OSWorld-Verified (83.0%), Terminal-Bench 2.1 (78.0%), GDM-MRCR v2 million token long context (54.0%), 3.6 Flash is generally better than 3.5 Flash and 3.1 Pro; but it is still inferior to DeepSWE GPT-5.6 Luna (67%) and Grok 4.5 (54%), SWE-Bench Pro is also lower than GPT-5.6 Luna’s 64.7%. The Elo for Knowledge Work GDPval-AA v2 increased from 1349 to 1421, but the absolute ranking is still in the middle. The general view in the industry is that Google still has advantages in speed and cost. 3.5 Flash-Lite and 3.6 Flash rank among the top two in the Artificial Analysis list in terms of output speed. However, in terms of intelligence and unit cost, they have been approached or even overtaken by models such as DeepSeek, MiniMax, GLM, and Grok. The real question lies on the delayed Pro flagship - once there are further setbacks in Gemini 4 pre-training, Google's voice in the model competition may be further under pressure.
-
The biggest controversy comes from the "reduced allocation" question. The official emphasis is on being more efficient, cheaper, and more token-saving, but a large number of actual tests in the community point to the regression of core capabilities. Added to the background of flagship delays, talent exodus, and market value pressure, the outside world will inevitably wonder whether Google is "patronizing cost savings and forgetting to mention IQ." The intelligence score of Artificial Analysis is flat, which further supports the statement that "cost reduction does not increase intelligence". In terms of security, 3.6 Flash has launched an upgraded version of Frontier Safety protection, focusing on two types of abuse: CBRN (chemical, biological, radioactive, nuclear) and network attacks. It has enhanced anti-jailbreak capabilities while minimizing false rejection of normal needs. The model card shows that text security has a slight regression compared to multi-language security 3.5 Flash (-1.35%, -5.45%), and the tone has also dropped slightly. However, Google said that manual review confirmed that most of the losses were false positives and were not serious. Known limitations still include hallucinations, occasional slowdowns or timeouts, and the knowledge deadline is March 2026.
-
3.6 Flash is suitable for teams where cost is the first constraint, workloads are mainly multi-step tool calls and long-term planning, and output tokens dominate billing: multi-modal tasks such as agent coding, computer operations, long document and very large code base reasoning, and mixed graphics, text, audio, video and PDF. It is also suitable as a "deputy" in high-throughput scenarios, such as batch feature extraction, web design generation, and translation summary receipts. If what you need is PhD-level scientific reasoning, the strongest multi-modal understanding, or the most powerful Google Brain you can buy today, then it cannot replace Gemini 3.1 Pro, let alone the 3.5 Pro, which is not yet GA. For projects that require high front-end visual accuracy, it is recommended to attach clear design specifications or directly use a model that is better at style. For users who have a preference for domestic models and value the ultimate cost-effectiveness, DeepSeek V4 Flash and Kimi K3 are also unavoidable comparison items.
-
Gemini 3.6 Flash is a workhorse that is more economical, faster, and more obedient. It has pushed Google's engineering efficiency and cost advantages in the Flash line to new heights, but it failed to bring surprises in terms of the upper limit of intelligence. The absence of the flagship Pro has magnified this gap. For enterprise-level agent workloads, it deserves the first regression test, but if you are looking forward to a capability jump, the answer will have to wait until Gemini 4.
User Reviews
-
JerryGray_2022—The free version is still available, let’s play around with the multi-modality first. -
Scott.HughesK16—My conclusion is to treat it as a work horse rather than a thinker: it is strong in agentic coding, computer operations, long documents and very large code base reasoning. Complex front-end vision and the strongest reasoning should be left to others; for enterprises to run tasks in batches, the token cost reduction is more practical than a few points on the benchmark. Before migrating, you should use your own real process for a round. -
NAngu_fi—In Artificial Analysis, the intelligence is completely on par with 3.5 Flash, which means there is no progress. -
COmor—Instead of giving a design draft, let it be used as the evaluation center website. The image can be produced in one minute. The main title, indicator cards, and button levels are all clear, and the color matching is unified. The degree of completion is not bad; but if it is really going to go online, the designer and front-end will definitely need to optimize another round of interaction and mobile details. -
EpsilonEarn2962025—Fast is really fast, stupid is really stupid. -
Noah_Fisher—To be honest, I am quite disappointed with this generation. The basic decoding occasionally fails, the Chinese word selection is sometimes outrageous, the quality of the generated pictures and videos does not match the instructions, and the memory is confused and the wrong tools are adjusted when doing complex tasks. The most uncomfortable thing is that you obviously spent less tokens, but you feel that the reliability of the answer has also decreased. It is true to save costs, but don’t save IQ as well. -
清风_14—Can this model be selected directly in AI Studio? Do I need to apply? -
SPowell007—I think saving tokens is not a free lunch, but a clear trade-off: a model that is trained to stop earlier will naturally suffer on tasks that require continuous visual reasoning and fine spatial layout, so the rollover of front-end, SVG, and game materials is structural, not an individual case; in turn, text, code review, document processing, and tool arrangement are upgraded very cleanly. The key is that you have to test your own most difficult tasks instead of looking at the list. -
Lincoln830—I ran a small game using Antigravity, and the wood grain was even better than the previous generation, while the marble was completely useless. -
truePauletteBischoff_pro—For the same seat arrangement question, 3.6 Flash only costs 2,338 tokens, and 3.5 Flash costs 3,707 tokens. This alone is real money for a team that runs agents every day. -
Frank_Bennett_66—Don't leave the front-end work to it, it will be the scene of the rollover. -
常婷_1—In the actual SQBench test of Dune, the 3.6 Flash scored 51.8 points, which is more than 11 points higher than the previous generation. It ranked fifth among the participating models, which shows that it is not smarter, but more agile. -
EugeneWatsonSr—I specially opened the deep thinking mode test, and the logical reasoning and constraint text generation are indeed significantly stronger. I got the 24-point calculation and the password questions I answered wrong before. However, the code generation is even worse. The 3D racing car cannot directly report errors. There are also logic bugs in the replica billboard. The front-end pop-up window lacks interaction, which is a typical side effect of overthinking. -
Samantha_Martin_Plus—Gemini 4 has started pre-training, but 3.5 Pro has not been released yet, which makes me anxious. -
Keith_Williams_77—The official said that it has launched an upgraded version of Frontier Safety, focusing on two types of abuse: CBRN and network attacks. It is more resistant to jailbreaking but tries not to accidentally damage normal needs. The security aspect is more stable than the previous generation. -
PMorgan_66—Compared with GPT-5.6 Luna and Grok 4.5, it is not significantly more expensive, but it is indeed a bit behind in terms of intelligence. The price-performance ratio is a bit embarrassing. -
JHill_99_460—The price has dropped and the efficiency has increased, but before the Pro was released, I always felt that Google was straining its computing power. -
Norman662—The 3.5 Flash-Lite from the same batch is even more powerful, with 350 tokens per second and an input cost of only $0.3. The high-throughput document processing can directly replace the work of the 3.6 Flash, which is really delicious. -
HawkHedge6902025—It was added to GitHub Copilot within a few hours of release. The third-party access speed is quite fast. Builders generally welcome the price reduction, but no one believes that Pro will really come on time this time. -
胡杰—It takes 9.4 seconds to answer the logic question, which is much faster than the 12.4 seconds of the 3.5 Flash. -
JackMurphy—Feedback from customers such as Hebbia and Harvey who are engaged in knowledge work is very positive, saying that it is outstanding in multi-modal tasks such as document parsing, chart analysis, and report drafting. Figma and JetBrains have also come out to support it. It seems that the office scene is much more reliable than writing front-end. If it really wants to be implemented in enterprise knowledge base, it will be a convenient role. -
bigrabbit586—The so-called knowledge ends in March 2026, but in actual testing, I basically don’t know anything about things after 2025. It actually stops at January 2025. This update is a bit confusing. -
SusanJohnsonQ82—Our team connected 3.6 Flash to the agent pipeline and ran it for two weeks. The most intuitive feeling is that the cost of single-task tokens has indeed come down. A long process has been reduced from an average of 2.7 minutes to 1.3 minutes, and tool calls have been reduced. However, for any work involving the front-end UI or space layout, the component nesting outputted by it is often messed up and the interaction logic is fragmented. In this scenario, we still fall back to Claude. The overall purpose is to use it as the main force to save money, not as an all-round brain. -
redostrich284—It is really stable for document parsing and long context retrieval, with 91.8% of the long text score there. It is suitable for feeding large code bases, but don’t count on front-end and spatial reasoning.