GLM-5.1

智谱AI发布的旗舰级开源大模型,全球首个实现8小时持续工作能力,编程能力超越Claude Opus 4.6

In-depth Report

  • GLM-5.1 is the flagship open source large model released by Zhipu AI on April 8, 2026, achieving a major breakthrough in domestic models in the field of software engineering. This model is the world's first open source model to verify its 8-hour continuous working capability through real engineering tasks. It achieved 58.4 points in the SWE-Bench Pro benchmark test, which is closest to real software development, surpassing GPT-5.4 and Claude Opus 4.6, and topped the list of global open source models for the first time. GLM-5.1 supports 200K ultra-long context, has excellent self-correction and long-term optimization capabilities, and can continue to work autonomously in hundreds to thousands of iterations, completing a complete closed loop from planning, execution to iterative optimization. In terms of pricing, the input price for the 0-32K context range is US$0.86/million Tokens, and the output price is US$3.5/million Tokens, which is about one-fifth to one-seventh of competing products.

  • GLM-5.1 is independently developed by Zhipu AI and released as open source under the MIT license. This model is optimized based on the 744B MoE architecture of GLM-5. It is the third model released by Zhipu in 2026. Previously, GLM-5 was released on February 12, 2025, and GLM-5-Turbo was released on March 16, 2025. In less than three months, Zhipu released three models in succession, forming a complete product matrix. Nine of the top ten Internet companies in China have connected to Zhipu AI's models. From the perspective of technological development, the core breakthrough of GLM-5.1 lies in its long-range task processing capabilities. Unlike traditional models, which are interactively called in minutes, GLM-5.1 can work continuously for a long time in a single task, covering hundreds to thousands of iterations. It can independently complete the complete closed loop of "experiment → analysis → adjustment → re-verification", repeatedly perform self-checks and path corrections at key decision-making nodes, and finally deliver engineering-level results. When GLM-5.1 was released on April 8, 2026, Zhipu simultaneously announced that the token price would increase by 10% again. After the release, Zhipu's stock price rose for three consecutive days, with its market value exceeding HK$400 billion, reflecting the market's recognition of its technical capabilities. Industry evaluation believes that "GLM-5.1 is not just a stronger model, but the opening of a new technical paradigm", marking the transformation of AI from "answering questions" to "completing projects".

  • As a flagship base model, GLM-5.1 has the following core features. In terms of thinking modes, the model provides a variety of thinking modes, covering different task requirements. Users can choose the appropriate way of thinking according to specific scenarios. In terms of output mode, GLM-5.1 supports real-time streaming response, which can significantly improve the user interaction experience and allow users to instantly see some results during the generation process. In terms of tool calling, GLM-5.1 has powerful Function Call capabilities, supports external tool integration and function calling, and can flexibly call external MCP tools and data sources to expand system capabilities. In terms of context processing, the model supports an intelligent caching mechanism to optimize performance in long conversation scenarios. It also supports output in structured formats such as JSON to facilitate subsequent processing. From the technical parameters, the input mode and output mode of GLM-5.1 are both text, the context window reaches 200K tokens, and the maximum output Tokens is 128K. These parameters enable GLM-5.1 to handle extremely long documents and complex multi-turn dialogue scenarios. In terms of recommended application scenarios, GLM-5.1 is particularly good at Agentic Coding scenarios. It is optimized for typical scenarios such as Claude Code and OpenClaw, and is suitable for real engineering tasks with multi-stage and strong dependencies. In addition, this model is also suitable for scenarios such as general dialogue, creative writing, Artifacts/front-end development, and Office productivity, and can complete the production tasks of complex documents such as PPT, Word, PDF, and Excel. Judging from the benchmark test performance, GLM-5.1 scored 58.4 points in the SWE-Bench Pro benchmark test, setting a new best performance in the world and surpassing GPT-5.4, Claude Opus 4.6 and Gemini 3.1 Pro. In the Terminal-Bench 2.0 and NL2Repo benchmark tests, GLM-5.1 also entered the top three in the world and ranked first among open source models. The comprehensive ability and coding ability are aligned with Claude Opus 4.6, and the programming ability reaches 94.6% of Claude Opus 4.6.

  • API pricing for GLM-5.1 varies based on context length. According to data from the 302.AI platform, the input price for the 0-32K tokens context range is $0.86/million Tokens, and the output price is $3.5/million Tokens; the input price for the 32K-200K tokens context range is $1.2/million Tokens, and the output price is $4/million Tokens. For large purchases, users can contact their account manager to enjoy exclusive discounts. Compared with the previous version GLM-5, the price of GLM-5.1 has increased. The input price of GLM-5 in the range of 0-32K tokens is 0.6 USD/million Tokens, and the output price is 2.6 USD/million Tokens. The price increase for GLM-5.1 is approximately 43% (input) and 35% (output). Compared with competing products, GLM-5.1's programming scenario pricing is close to the level of Anthropic's Claude Sonnet, but its overall cost-effectiveness still has significant advantages. According to user reviews, the programming capability of GLM-5.1 reaches 94.6% of that of Claude Opus 4.6, while the price is only about 1/5 to 1/7 of competing products. It is evaluated by users as "using 30% of the money to achieve 94% of the capabilities." When GLM-5.1 was released on April 8, 2026, Zhipu simultaneously announced that the token price would increase by 10% again. After the release, Zhipu’s stock price rose for three consecutive days, with its market value exceeding HK$400 billion. This market reaction shows investors' recognition of GLM-5.1's technical capabilities and their confidence in the future development of Zhipu.

  • Judging from the positive reviews, users generally believe that the programming ability of GLM-5.1 has reached the first echelon level of Claude Opus 4.6 and can independently complete complex engineering tasks. Actual test cases show that GLM-5.1 can completely derive multiple solutions in logical reasoning tasks, and perform self-verification and path correction at key nodes. As the world's first open source model to verify its 8-hour continuous working capability through real engineering tasks, GLM-5.1's performance in long-term tasks has been recognized by users. Its self-correction and long-term optimization capabilities enable it to continuously maintain the consistency of task goals across hundreds of rounds and thousands of tool calls. User feedback shows that GLM-5.1 tends to complete things and actively fill in the details. Compared with competing products, it is better at organizing the overall structure and completing details of project-level tasks. In terms of cost performance, users commented that "30% of the money is used to achieve 94% of the capabilities." GLM-5.1 has an obvious price advantage and has high practical value for handling 90% of standard tasks. In the front-end development test, the web page portfolio generated by GLM-5.1 has a complete interactive system, scrolling parallax effect, mouse following soft light effects, etc. The overall vision is "restrained and natural" and received high evaluation. Judging from the negative feedback and improvement suggestions, some users reported that the official channel has problems such as difficulty in obtaining subscriptions, lags and delays in use, and the computing power supply cannot fully meet user needs. This has led some users to turn to third-party deployment channels for solutions. Some users pointed out that GLM-5.1 is not good at "subtraction" and is slightly weaker in expressing the sense of extreme design. It is more like an engineering model with strong execution rather than pursuing single-dimensional limits. In tasks such as complex animation, GLM-5.1 generates a large amount of code, which may cause certain performance overhead. Some users reported slower response times when the tasks were complex, but the overall ability was affirmed.

  • The release of GLM-5.1 has attracted widespread attention in the industry and is regarded as an important milestone in the development of domestic large models. In the field of open source, GLM-5.1 is the first echelon of programming capabilities in large open source models in the world. It surpassed GPT-5.4 and Claude Opus 4.6 in authoritative evaluations and topped the list of global open source. It is called the "Claude Opus" of the open source industry by the industry. This is a milestone breakthrough for domestic large-scale models in the field of software engineering. For the first time, it has reached the international top closed-source model level in terms of core engineering indicators. Industry evaluation believes that GLM-5.1 is not just a stronger model, but the opening of a new technical paradigm, marking the transformation of AI from "answering questions" to "completing projects." Judging from the market response, in less than three months, nine of China's top ten Internet companies have rushed to access Zhipu AI's model, forming widespread recognition in the industry. A large number of companies have announced on social media and official websites that they have "connected", including Internet companies, cloud service providers, software manufacturers, chip companies, large, medium and small. After releasing GLM-5.1 and announcing a price increase, Zhipu's stock price rose for three consecutive days, with its market value exceeding HK$400 billion, reflecting investors' recognition of its technical capabilities. The successful release of GLM-5.1 has pushed China's large-scale models from the "catch-up" stage to the "tackling" stage, and the industry competition landscape has changed. In terms of the competitive landscape, GLM-5.1’s main competitors include Anthropic’s Claude series, OpenAI’s GPT series, and Dark Side of the Moon’s Kimi series. In April 2026, Kimi K2.6 and GLM-5.1 were released one after another, both emphasizing long-cycle coding and agent capabilities, but the technical routes are completely different. GLM-5.1 focuses more on engineering delivery capabilities, while Kimi K2.6 has its own focus on multi-modal understanding.

  • Judging from user feedback, the main controversy over GLM-5.1 currently focuses on the supply of computing power. Some users have reported that official channels have problems such as difficulty in obtaining subscriptions, lags and delays in use. The tight supply of computing power is the main problem currently. This reflects that Zhipu needs to further increase its investment in computing infrastructure to meet the rapidly growing user needs. From a technical perspective, some users pointed out that GLM-5.1 is more biased toward engineering models, is not good at "subtraction", and is slightly weaker in expressing a sense of extreme design. This means that GLM-5.1 is more suitable for project-level tasks that require complete delivery, rather than scenarios that pursue a single point of ultimate performance. In addition, GLM-5.1 announced a price increase shortly after its release. Although the market responded positively, it also caused some users to worry about cost control. For enterprise users who need to use GLM-5.1 on a large scale, the increase in API costs needs to be taken into consideration in the overall budget.

  • GLM-5.1 is particularly suitable for the following user groups: users with complex programming tasks, especially real engineering tasks that require multi-stage and strong dependencies; users who require long-term software engineering delivery, and can continuously work for 8 hours to complete the complete project delivery; projects that require continuous iterative optimization, with self-verification and path correction capabilities; cost-sensitive projects, user reviews show that its price is only 1/5 to 1/7 of competing products. For the following scenarios, GLM-5.1 may not be the best choice: Scenarios that pursue the ultimate performance in one dimension are more inclined to engineering models; scenarios with a sense of extreme design are not good at "subtraction" and their expression is slightly weaker. From an alternative perspective, if users need to handle the most difficult 10% of tasks, they may still need to use cutting-edge models such as Claude Opus 4.6 or GPT-5.4; but for the remaining 90% of standard tasks, GLM-5.1 has high practical value.

  • GLM-5.1 is the flagship open source large model launched by Zhipu AI, which has achieved a major technological breakthrough in the field of software engineering. Its core advantages include: 8-hour long-distance task processing capabilities, SWE-Bench Pro's world-leading open source programming capabilities, 200K ultra-long context support, and significant price advantages compared to the world's top closed-source models. As the world's first open source model that verifies its 8-hour continuous working capability through real engineering tasks, GLM-5.1 redefines the transformation of AI from "answering questions" to "completing projects." Its self-correction and long-term optimization capabilities enable it to continuously maintain the consistency of task goals and deliver engineering-level results over hundreds of rounds and thousands of tool calls. Judging from market performance, nine of China's top ten Internet companies have connected in less than three months, and the market value has exceeded 400 billion Hong Kong dollars, which proves that GLM-5.1's technical strength has been recognized by the industry. Looking to the future, the successful release of GLM-5.1 marks that China's large models have entered the "attack" stage from the "catch-up" stage, and the status of domestic large models in the global AI competition will be further enhanced.

User Reviews

  • 头像
    NAdams_Pro
    睡前把需求丢给它,早上起来活儿真干完了,中间自己规划自己debug,跨几十步还记得最初的约束,这个长程任务能力国产里确实是断档第一。

  • 头像
    Logan.Cox
    能力没得说,就是慢到离谱,一个复杂任务跑了一个多小时,急起来真想砸键盘。

  • 头像
    SeanGray_7
    分享个真实案例。我之前一直用 Typeless 那个 mac 语音输入,年费一千块,主要拿来做 vibe coding。前几天看到 GitHub 上有人发了一段特别详细的提示词,完整描述了一个菜单栏语音输入 app 的需求,我直接原封不动扔给 GLM-5.1 跑。它自己拆模块、写 Swift、遇到编译冲突自己定位改掉,全程没问我一句,大概二十分钟就吐了个带 Makefile 的完整项目,build 出来签名好的 app 直接能跑,按住 Fn 说话底部弹胶囊悬浮窗,波形跟着声音跳,松手文字准确填进光标位置。整个过程只用了五小时额度里不到百分之十。这成品覆盖了 Typeless 九成以上功能,代码还全在我手里,那一千块我大概率不续了。

  • 头像
    Vincent_GonzalezX
    看到有人测睡一觉让它从零搭 Linux 桌面,八小时一千两百多步,早上起来窗口管理器、状态栏、应用、VPN 管理器全齐了,说是相当于四人团队一周的量,这画面还挺科幻的。

  • 头像
    何莉霞
    国产之光实至名归了这次。

  • 头像
    Alan.Thompson_Max
    冷静点说,官方宣传 Pro 套餐额度是 Claude Pro 的十五倍,我实测完全是虚标。跑完一个复杂任务 weekly usage 直接来到百分之八,半天用下来就百分之十了,高强度用根本不够。编码质量确实没啥大毛病,但速度慢加用量虚标这两个硬伤摆在那,所谓性价比优势其实没想象中那么香。

  • 头像
    AltSeasonHughes
    200K 上下文也好意思当亮点吹?人家 Kimi 256K、DeepSeek 都 1M 了,整个代码库塞进去还是有点吃力。

  • 头像
    Megan.Wright_6638
    最戳我的一点是遇到坑不喊人。让它做本地记账应用,装 better-sqlite3 那个包要编译,我环境里没 C++ 工具链,换别的模型这里肯定停下来让我先装,结果它自己发现编译失败,直接改用 sql.js 纯 JS 方案接着往下跑,最后浏览器直接能用。

  • 头像
    JCarter_202105
    作为重度 coding 用户认真评一下。SWE-bench Pro 上它刷新了全球最佳,真实 GitHub 仓库定位修 bug 这种最硬的指标能压过 GPT-5.4 和 Opus 4.6,代码三项综合全球第三国产第一。我印象最深的是那个向量数据库优化的例子,655 轮迭代自己从全库扫描切到 IVF 分桶、加半精度压缩、量化粗排、两级路由,硬是把吞吐从 3108 QPS 推到 21472,六点九倍。这已经不是代码生成器了,是会自己找瓶颈换策略的优化器。

  • 头像
    HhshNet
    架构设计和 UI 审美还是差点意思,得配脚手架。

  • 头像
    Abigail_Turner_99
    真香!用它做了个后台管理系统,从数据库设计到接口实现全程自动,跑通了!之前用其他模型经常卡壳,这个8小时连续工作不是吹的,确实稳。

  • 头像
    邹玉
    用了两周GLM-5.1写后端代码,性价比真的没话说!日常CRUD和小功能开发完全够用,比Claude便宜太多了,终于可以放开膀子写代码了!

  • 头像
    N_athanBaker
    国内抢购这操作是真的服了,定闹钟十点开售照样秒没,钱都得抢着交。最后走国际版年付才买到,也是无语。

  • 头像
    Judith.Myers_2023
    强烈推荐给预算有限的团队!用它替代Claude处理日常需求,一个月上千的订阅费直接砍掉大半,省下来的钱可以买点别的工具。

  • 头像
    PWard_Max
    说实话有点超出预期。本来以为国产模型就那样,结果GLM-5.1的长程任务能力真不错,测试用例自己就能写完整,还知道回头校验错误。

  • 头像
    VYW9X1F
    吐槽一下高峰期涨价的问题!下午三四点那会儿消耗额度直接翻三倍,用起来心疼。建议错峰使用,早上九点前或者晚上十点后性价比最高。

  • 头像
    C_Lree
    刚用GLM-5.1重构了一个半废弃的Vue项目,整体体验还行。不过复杂动画场景还是要手调,生成的代码量有点大,性能开销需要优化。

  • 头像
    Theresa_CookK
    国产之光名不虚传!用它做了个数据可视化大屏,从图表选型到交互逻辑全程AI辅助,最后交付的东西客户还挺满意的。

  • 头像
    ARussell520
    响应速度确实比Claude慢,但胜在便宜啊!用它处理90%的标准任务完全没问题,那10%的复杂场景再用Claude兜底,这样搭配最划算。

  • 头像
    r2ke2u1wtp
    前端调试起来稍微有点头疼,CSS样式纠缠不清的情况还是有的。不过后端代码质量挺高,接口设计和异常处理都比较到位。

  • 头像
    PamelaScott
    用GLM-5.1开发了一个小工具箱,涵盖七八个实用功能,累计跑了四十多个小时。中间偶发卡顿,但整体稳定性能接受。

  • 头像
    ADtur
    客观说,复杂多文件重构和架构设计还是Claude强一些。但日常开发用它足够了,省钱才是硬道理啊!

  • 头像
    潘悦勇
    长上下文处理确实强,10万行代码库丢进去分析,它还记得前面说过的约束,不会突然失忆。这个能力对大型项目帮助很大。

  • 头像
    TGomez_9915
    订阅确实难抢,每次放号都要蹲点。希望智谱能扩大算力供应,需求太大了根本不够用。

  • 头像
    Natalie.Hicks168
    用它写了几个Python脚本,处理Excel和数据清洗太方便了!批量操作和格式转换都能自动完成,省了不少重复劳动。

  • 头像
    苏月洋
    说实话,界面可以再简洁一些。不过功能是真的全,用起来还挺顺手的。

  • 头像
    orangecat977
    配合Claude Code使用效果更佳!用它规划任务和生成代码框架,Claude负责最终优化和调试,分工明确效率翻倍。

  • 头像
    NicholasGutierrez_X
    生成代码的速度挺快的,就是等待的时候有点煎熬。建议智谱优化一下流式输出的体验。

  • 头像
    iJoséLozano
    用了大概一个月,感觉进步很明显。之前GLM-5的一些问题都修复了,死循环的情况基本没再遇到过。

  • 头像
    Shirley.Moore007
    用它做了个项目管理系统,包含用户、订单、权限等模块。从数据库设计到前端页面全程AI辅助,两天就交付了MVP版本。