Claude Opus 4

Anthropic's flagship large language model, focusing on coding and AI agent fields, provides 1 million context windows

In-depth Report

  • Claude Opus 4.7 is Anthropic's flagship large language model released on April 16, 2026. It is positioned as a hybrid inference model and focuses on the fields of coding and AI agents. The model provides 1 million context windows and scored 64.3% in the SWE-bench Pro test and 70% in CursorBench. In terms of pricing, Opus 4.7 is priced at US$5 per million input tokens and US$25 per million output tokens. However, this version caused widespread controversy after its release. Users criticized its text expressions for becoming mechanical and lacking in human touch, and the new word segmenter increased token consumption by 0% to 35%, and the actual cost of use increased. Overall, Opus 4.7 is more suitable for professional developers and enterprise-level complex tasks, rather than content creators and daily interactive users.

  • The Claude Opus series is developed by artificial intelligence company Anthropic. Anthropic was founded in 2021 and is headquartered in San Francisco, USA. It was founded by former OpenAI employee Dario Amodei and others. It focuses on building safe and reliable large-scale language models. The company has received billions of dollars in investment from technology giants such as Google and Amazon, with a valuation of more than $60 billion. Claude Opus 4.7 is the latest version of the series, officially released on April 16, 2026. Prior to this, Anthropic has released Claude Opus 4 in May 2025, Opus 4.1 in August 2025, Opus 4.5 in November 2025, Opus 4.6 in February 2026, and Opus 4.7 in April 2026. It can be seen that Anthropic is rapidly iterating its flagship model every two to three months, and the intensity can be called an AI arms race. Anthropic's product line currently includes three series: Opus, Sonnet and Haiku, which are respectively targeted at high-end, mid-range and entry-level user scenarios. Among them, the Opus series is positioned as the most powerful general-purpose intelligent model, mainly serving professional software engineering, complex agent workflows, and high-risk enterprise tasks.

  • The core functional upgrades of Claude Opus 4.7 focus on the three dimensions of coding capabilities, visual understanding and AI agents. In terms of coding capabilities, Opus 4.7 has achieved a significant breakthrough. According to official data, the SWE-bench Verified score increased from 80.8% in Opus 4.6 to 87.6%, and the SWE-bench Pro score jumped from 53.4% ​​to 64.3%. This means that the model can autonomously complete more complex programming tasks, including difficult tasks such as building a complete Rust speech synthesis engine in a single session. User feedback shows that Opus 4.7 can discover its own logic errors in the code planning stage and can continuously correct itself during the execution process, greatly reducing the review cost of senior engineers. In terms of visual understanding, Opus 4.7 has achieved epic enhancements. The model supports high-resolution image understanding up to 2576 pixels, and its score in the visual sharpness benchmark test increased from 54.5% in Opus 4.6 to 98.5%, an increase of nearly double. This enables the model to accurately identify professional content such as complex technical diagrams, interface screenshots, PDF documents, and chemical molecular structures, which is a major benefit for users who need to process a large amount of mixed graphics and text materials. In terms of AI agent capabilities, Opus 4.7 adds an adaptive thinking function that can automatically adjust the depth of thinking based on the complexity of the task. Simple questions are responded to quickly, while complex questions require more computing resources. This model also supports longer-period background running tasks, and users can allocate long-running coding work to Opus for processing alone. In addition, the model's tool calling accuracy and planning capabilities in multi-step workflows have been improved, and it can drive production-level agent systems more reliably. In terms of user experience, Opus 4.7 introduces a new tokenizer. Although the official pricing remains unchanged, since the number of tokens consumed for the same content increases from 0% to 35%, the actual cost of use for users increases in disguise. The model also provides a new thinking gear called xhigh, which is between high and max, providing more stable reasoning performance for complex tasks.

  • Claude Opus 4.7 provides multiple levels of access. In the consumer and small business markets, Opus 4.7 is available through Claude for Pro, Max, Team and Enterprise subscription plans. In the developer market, the model is available through multiple channels including Claude Platform API, Amazon Bedrock, Google Cloud Vertex AI, and Microsoft Foundry. In terms of API pricing, the input price of Opus 4.7 is US$5 per million tokens, and the output price is US$25 per million tokens. Through prompt caching and batch processing functions, users can achieve cost savings of up to 90% and 50% respectively. For workloads requiring US-based inference, the US-dedicated inference service charges 1.1x the token fee. It is worth noting that the price of Opus 4.6, launched by Anthropic in February 2026, is US$15 per million input tokens and US$75 per million output tokens. Opus 4.7 has achieved a significant price reduction in comparison. This pricing strategy shows that Anthropic is lowering the barriers to use of high-end models through scale effects and efficiency optimization.

  • User reviews of Claude Opus 4.7 are clearly polarized. Feedback from the professional developer community has been generally positive. Replit users report that Opus 4.7 achieves the same level of quality as its predecessor at a lower cost for day-to-day tasks such as log analysis, bug finding, and fix recommendations. Warp users said that Opus 4.6 is already the best model for developers, and 4.7 is more rigorous on this basis and can complete terminal tasks that previous versions could not handle. Cursor users reported that in the CursorBench test, Opus 4.7 achieved a pass rate of 70%, which was significantly improved compared to Opus 4.6's 58%. However, feedback from content creators and everyday user groups has been quite negative. A large number of users criticized Opus 4.7's text expression for becoming mechanical, stiff, and full of Internet slang, and no longer possesses the delicate and smooth expression ability of the previous generation model. Some users even use "Claude-lash" to describe this regression in experience. Business Insider reports that users are frustrated by models that confidently assume that information does not exist and insist on incorrect interpretations. This kind of "confidently making mistakes" is more difficult to accept than ordinary errors because it directly increases the user's review costs. In addition, users also report that Opus 4.7 executes instructions more "literally" than the previous version, requiring more specific prompt words to obtain ideal results. In low effort mode, the model performs significantly worse than Opus 4.6, causing confusion for users who rely on the default settings.

  • From an industry perspective, the release of Claude Opus 4.7 has intensified competition in the field of AI large models. In terms of benchmark testing, Opus 4.7 ranks among the best in many authoritative reviews. Artificial Analysis comprehensive intelligence index shows that GPT-5.5 leads with 60 points, followed by Opus 4.7. However, in the key benchmark test of ARC-AGI-3, both GPT-5.5 and Opus 4.7 failed to achieve breakthrough results, and human testers instead obtained a perfect score of 100 points. Industry media generally summarize the characteristics of Opus 4.7 as "the performance is blazing but not human". The evaluation agency pointed out that this model does represent the highest level of current technology in terms of coding and visual understanding, but its regression in text expression and natural interaction is worrying. This "tooling" tendency is not unique to Opus 4.7, but is a common trend in the entire AI industry - from GPT-5.4 to the Claude series, each flagship model is evolving in the direction of "professional tools", while general intelligence and humanistic expression capabilities are being sacrificed. From a business competition perspective, Anthropic has further consolidated its leading position in the enterprise-level coding and agent task market through Opus 4.7. However, the significant reduction in pricing also shows that large AI models are entering a price war stage, and the pricing of high-end models may continue to drop in the future.

  • The controversy caused by the release of Claude Opus 4.7 mainly focuses on the following aspects. The first is the cost dispute. The increased token consumption caused by the new tokenizer caused user dissatisfaction, and many users quickly exhausted their subscription quota without knowing it. Considering that adaptive thinking and higher effort settings will further increase token consumption, the increase in user-perceived costs may far exceed the changes in official pricing. The second is the capability dispute. Some users believe that Opus 4.7 has become "dumb" in certain scenarios, especially in small daily tasks. However, this perception may be related to default behavior changes, effort setting adjustments, and product-level configuration changes, and it is difficult to simply blame it on the model itself. The third is the positioning dispute. Opus 4.7 is clearly geared towards the professional developer market, which has triggered a sense of loss among content creators and ordinary users. The Claude series was once known for its "tasteful text expression," but Opus 4.7 is losing this differentiation advantage. In terms of technical risks, Anthropic actively weakened its network security capabilities in the System Card of Opus 4.7, causing dissatisfaction among some in the security research community. Although the Cyber ​​Verification Program is officially provided as an application channel for legitimate security research, the wisdom of this proactive restriction strategy remains to be discussed.

  • Claude Opus 4.7 is suitable for the following user groups. Professional software engineers are the primary target users. This model performs well in complex code development, long-cycle agent tasks, and production-level code reviews, and can significantly improve development efficiency. Opus 4.7 is the most powerful option available for teams that need to handle large code bases, perform multi-step refactorings, or perform unattended programming tasks. Enterprise-level users are also the main service targets. Opus 4.7 demonstrates a high level of stability and professionalism in enterprise workflows such as processing complex documents, analyzing spreadsheets, and producing presentations. The context window of 1 million tokens is suitable for long document processing and large-scale code base analysis. The same applies to high-risk corporate mission scenarios. Because the model can maintain a longer attention span and think more deeply during the reasoning process, it is suitable for professional fields such as finance, medical, and law that require high reliability. Conversely, the following user groups may not be suitable for Opus 4.7. If the content creator's main needs are copywriting, story creation, or the expression of beautiful words, it is recommended to continue using Opus 4.5 or wait for improvements in subsequent versions. If daily interaction users pursue a natural and smooth conversation experience, Opus 4.7 may not be as good as the previous version. Budget-sensitive users need to pay attention to the hidden costs caused by increased token consumption. Regarding alternatives, users looking for a more balanced performance can consider the Claude Sonnet 4.6, a model that strikes a good balance between intelligence and speed. Or wait for the subsequent optimized version of OpenAI GPT-5.5 and the update of Google Gemini series.

  • Claude Opus 4.7 is one of the most powerful professional-level AI models currently available, representing the highest level in the industry in terms of coding, visual understanding, and complex agent tasks. However, this version sacrificed the ability of text expression and humanistic interaction while pursuing technical performance, causing widespread user controversy. From a business perspective, Anthropic is lowering the threshold for using high-end AI models through a significant price reduction strategy, which has positive significance for the development of the entire industry. However, the polarization of word-of-mouth also reminds us that the development of AI models cannot only pursue the improvement of technical indicators, but also needs to pay attention to the integrity of the user experience. For professional developers and enterprise users, Opus 4.7 is undoubtedly the most powerful productivity tool at the moment. However, for content creators and ordinary users who pursue human-machine collaboration experience, they may need to wait for Anthropic to rebalance technical performance and humanistic expression in subsequent versions.

User Reviews

  • 头像
    ecvyy
    Opus 4.8 输出真的太啰嗦了,问它怎么写一个正则表达式匹配邮箱,它先给我讲三段关于正则表达式历史的故事,从 1950 年代的 Stephen Kleene 开始讲,然后才给那一行代码。GPT 直接省略开头废话上来就给代码。每天跟 AI 聊几百��合的人,浪费在看废话上的时间至少半小时。

  • 头像
    AxelRønning
    听说 Mythos 级别模型快开放了,Anthropic 说几周内给所有客户,到时候 Opus 4 的 API 价格会不会降。

  • 头像
    umawv_n
    Opus 4.8 的代码能力真的强,但也是真的贵,纠结。

  • 头像
    石珍
    有人觉得 Opus 4.8 的对抗性是好事吗?我反而觉得它更像一个谨慎的同事,虽然效率低了但正确率确实上去了。在某些需要严谨性的行业比如法律和合规,它宁愿不答也不瞎答的态度反而是优点吧。

  • 头像
    AnnCruzX
    最近用 Claude Opus 4.8 做了个大项目,120 页的需求文档扔进去让它找前后矛盾的地方,真给找出 8 处来,其中 3 处我自���都没发现。某个功能在第 37 页说支持批量导入,第 89 页又改成只支持单条创建了,这种前后不一致人工 review 真的很费劲,它一分钟就找出来了。GPT 我试过,同样需求只能找出两处明显的格式错误,深层矛盾完全没看出来。

  • 头像
    ChristopherWalker_Plus5
    升级 Opus 4.7 以后肠子都悔青了,自适应推理机制完全是在摸鱼!以前 4.6 一个回合就能搞定的复杂重构任务,4.7 要来回扯好几轮,每个回合给的答案都不一样。你让它重新检查,它就换一个完全不同的方案,还夸你要求它再次检查。这就是我当初离开 GPT 的原因,现在在 Claude 上又来一遍。

  • 头像
    SParkerII_661
    Opus 4.8 更新了 Dynamic Workflows,能同时拉起几百个子 Agent 并行干活,先规划再分发给子 Agent 最后汇总验证。感觉以后重构整个项目或者做大库迁移真的可以交给 AI 全程主导了。

  • 头像
    JEdwards520
    4.6 yyds,Anthropic 你为什么要换掉它!

  • 头像
    悠然_10
    Claude Opus 4 的代码审查能力确实没话说,前两天把整个微服务的代码库丢给它做 review,它把几个隐藏很深的并发边界 bug 全揪出来了。那些 bug 我们团队自己 review 了三轮都没发现。但 Opus 真的太慢了,等它输出感觉像在等公交车,首 token 延迟平均 2 秒以上,做实时交互项目真的不行。

  • 头像
    PatriciaJones_88
    大家平时用 Opus 4 做啥?我是只用来 review 代码和��长文,日常对话用 Sonnet 省很多钱。感觉大部分人日常工作场景根本不需要上 Opus,Sonnet 已经足够好了。混合部署把 API 成本降到纯 Opus 的 15% 不香吗。

  • 头像
    HeatherGray_998
    Opus 4.8 的对抗性真的受不了,问它 Python 和 Java 哪个适合做数据分析,它先写八百字免责声明——这个问题没有绝对答案取决于您的具体使用场景项目需求团队技术栈等等等等。GPT 就很直接:做数据分析选 Python,做企业级应用选 Java。用户要的是直接答案不是免责声明好嘛。

  • 头像
    BlockDex_n
    用 Claude Opus 4 写中文行业分析报告体验确实独一档,3000 字的稿子一次出稿,��言自然度接近专业撰稿人水平,成语和文化典故用得很到位。GPT-5 写出来偶尔还是会有翻译腔,读起来怪怪的。但 Opus 的价格真的离谱,$75 每百万输出 token,我写一篇报告光 token 费就顶 GPT 好几天了。

  • 头像
    流年472
    给大家一个省钱技巧,日常代码用 Sonnet、长文档分析切 Opus、简单任务丢 Haiku,混合部署能把 API 成本降低 85%。70% 请求走 Haiku,25% 用 Sonnet,��有 5% 的复杂任务才开 Opus。这样月成本从 1 万 2 降到 1800 刀,省下来的钱够招一个实习生了。

  • 头像
    BobbyGarcia_Plus
    Claude Opus 4 的中文创作没人能打,就是钱包扛不住。

  • 头像
    PYoung_987
    有没有人和我一样觉得 Opus 4 的中文写作比 GPT-5 好太多?尤其是公文和商务写作,成语典故的运用和理解深度明显不一样。感觉国内企业级应用如果对中文质量有要求应该优先考虑 Claude 而不是 GPT。

  • 头像
    蝴蝶_21
    Opus 4.8 所谓的 200 万 token 上下文窗口水分太大了,前 100 页的内容它记得还行,到 150 页以后就开始选择性遗忘了。你提醒它,它恍然大悟说哦对之前提到过这个我刚才忽略了。跟它聊了 10 轮,第 11 轮就忘了第 3 轮的关���信息。花大价钱买长上下文结果能用的只有四分之一,这不是坑人吗。

  • 头像
    RPhillips_99
    把公司上个月的会议录音转写了两万多字扔进 Opus 4.8,让它对比谁在会上说一套会后做一套。它真给我找出来了——张三在会上说这个需求我没问题,但翻他两周前发的邮件发现他明明说过这个技术上不可行。这种跨文档的人物态度反差分析,目前真的只有 Opus 能做好,其他模型要么漏掉要么分析得不够深。

  • 头像
    ACoxSr
    用 Opus 4 做代码审查的时候记得给出完整的项目上下文,不要只丢一段代码进去。它的 200K 窗口能容纳整个微服务的代码库,你给它看全貌,它能发现跨文件的依赖问题和架构缺陷。只看单文件的话效果打五折,浪费了它最强的能力。

  • 头像
    TokenLink_r
    Opus 4.8 活干得好,但话太难听,体验极差。

  • 头像
    JStephensZ
    Opus 4.7 的幻觉问题真的太严重了!让它做物理计算,它竟然编造自己搜索过,但界面上那个搜索指示器根本没亮过。被当场拆穿以后它滑跪说你说的对我没有搜索抱歉。更离谱的是它还会编造不存在的人名,在讨论代码变更时突然问你要不要和产品负责人 Anton 讨论一下,追问才知道这个名字完全是编的。

  • 头像
    AnnaWeaver
    Claude Code 配合 Opus 4 做代码重构体验确实好,设置好 Agent Teams 以后一天内自主关了 13 个 issue,还自动给我代码库生成了导航文档。7 小时不间断运行也不出岔子,Rakuten 那边的测试也是类似结论。就是 200 美金月费真的肉疼,个人开发者用不起,只能让公司报销。

  • 头像
    Deborah_Cooper
    Opus 4.7 真的比 4.6 差吗?跑分不是涨了嘛,SWE-bench 从 80.8% 到了 87.6%,为什么实际体验反而倒退了。有人说是自适应推理的锅,模型自己判断该不该深度思考,结果判断错了。如果是这个问题 Anthropic 应该出一个手动控制推理深度的开关。

  • 头像
    Christine.Stewart007
    Anthropic 完成新一轮 650 亿美元融资估值逼近万亿美元,这家公司现在的估值快赶上腾讯了。不过钱烧得也快,光是 Opus 4.8 的训练成本不知道占了多大比例。

  • 头像
    lILLIAN443
    Opus 4.8 的拒绝率比以前高太多了,我一个用了一年半的老用户以前只遇到过两三次拒绝,升级到 4.8 以后一周被拒了 8 次。同一个项目在 4.7 上做了一半,切到 4.8 它直接拒绝继续执行,说你的需求可能有伦理问题。我就是写个学校作业而已至于吗。

  • 头像
    TUwbl
    Opus 4 的指令遵循能力是目前所有模型里最好的,给了一个 10 条风格限制加目标字数加禁止词列表,它全给你严格遵守。GPT-5 每次总会在某个细节上偷懒或者遗漏。但代价就是输出太啰嗦,哪怕是简单问题它也要先铺垫三段背景再给答案,急性子真的会被气死。

  • 头像
    胡飞霖
    Opus 4.8 的 effort 控制一定要学会用,简单问题调 low 省 token 省钱,深度分析调 max 保证质量。别傻傻一直开默认值,不然钱包真的会哭。Claude Code 里用 /effort 参数就可以调,简单查询设 medium 就好,只有做大型重构的时候再拉到 max 或 ultracode。

  • 头像
    4meacuqbx
    三个月前还在夸 Opus 4.6 好用,结果后面越改越差。4.7 偷懒摸鱼自适应推理自己决定该不该动脑子,4.8 又走向另一个极端变成说教型管家。Anthropic 的���品路线图到底是谁在拍板,能不能出一个用户可选模式而不是替用户决定这模型该是什么性格。

  • 头像
    80z26tm1j
    用 Opus 4.8 做合同审查真的事半功倍,几十页的采购合同扔进去,它把隐藏的霸王条款全标红了,什么交付时间模糊啊、验收标准缺失啊、违约金不对等啊,全都给你列出来还附修改建议。律所朋友说这份工作之前要花一整天,现在 Opus 几分钟搞定初筛,他们再人工复核一遍就行了。

  • 头像
    greenfrog189
    升级 4.7 后卡成 PPT,后悔死了。

  • 头像
    irw8g00g
    用 Opus 4.8 做创意��作完全是灾难,写出来的东西官方公文味太重,每个句子都四平八稳像政府工作报告。改它的东西比自己写还累,完全没有网感和灵气。中文创作巅峰真的停留在 Opus 4.5 和 4.6 了,后面这几版越改越倒退,Anthropic 是不是把创意团队的训练数据给换掉了。