Kimi K3 Open-Source Edition

Moonshot AI's largest open-source LLM in the world: 2.8 trillion parameters, native multimodality and a 1M-token context window, topping global coding benchmarks

In-depth Report

  • On July 16, 2026, Beijing Dark Side of the Moon Technology Co., Ltd. released Kimi K3, a large open source model with a total parameter scale of 2.8 trillion. This is currently the open source model with the largest parameters in the world. It topped the Frontend Code Arena front-end programming list with 1679 points, surpassing Claude Fable 5 and GPT-5.6 Sol. K3 is built based on the self-developed KDA hybrid linear attention mechanism and attention residual technology. It natively supports visual understanding and has a 1 million token context window. It is specially optimized for complex task scenarios such as long-range programming, knowledge work, in-depth research and multi-modal understanding. It shows cutting-edge performance in endurance programming and agent tasks, but its overall comprehensive intelligence still lags behind the strongest closed-source models.

  • Moonshot AI was founded in 2023 and is headquartered in Beijing. It is an AI startup company invested by Alibaba and Hong Kong Investment Management Co., Ltd. (Hong Kong Investment). The founder, Yang Zhilin, was previously an assistant professor at the Institute of Cross-Information at Tsinghua University. The core members of the team come from top universities such as Tsinghua and Peking University and technology companies such as Google and Meta. Kimi K3 has been in development for over a year and is the official successor to Kimi K2 (released in July 2025). In the meantime, iterative versions such as K2.5, K2.6, and K2.7 Code were also released. The release of K3 not only achieves a qualitative leap in parameter scale - jumping from K2's trillion level to 2.8 trillion - but also makes three key innovations in the underlying architecture: KDA hybrid linear attention mechanism (Kimi Delta Attention), attention residual technology (Attention Residuals) and Moon Clip second-order optimizer. It is worth noting that the release of Kimi K3 also caused an industry storm. Michael Kratzios, director of the White House Office of Science and Technology Policy, accused Dark Side of the Moon of developing K3 by distilling Anthropic's Fable model on social media, and threatened sanctions and an entity list. Dean Ball, OpenAI's head of strategic future, also publicly criticized it, calling it a "decelerationist force." Huang Zhenxin, head of the Dark Side of the Moon Enterprise business, responded clearly that the core of K3’s performance jump comes from the underlying original architectural innovation and is not a distilled copy of the existing model. Many authoritative media and AI research scholars also publicly stated that the accusation lacked technical basis.

  • Kimi K3’s core capabilities focus on the following directions. The first is programming ability, which is its most promising area. On the Frontend Code Arena front-end code list, K3 ranks first with 1679 points, surpassing Claude Fable 5 (1631 points) and GPT-5.6 Sol (1618 points). The actual programming experience is divided into two scenarios: one is used through the Kimi Code terminal tool, which supports /model command switching, and is suitable for long-term continuous programming tasks; the other is called through the Kimi API, which is compatible with the OpenAI SDK and is priced at $0.3 per million tokens for cache hits, $3 for input misses, and $15 for output. The second is long-distance task processing capabilities. Officially demonstrated cases include independently completing chip design and optimization for 48 consecutive hours, and reproducing the I-Love-Q universal relationship in the field of computational astrophysics in about two hours. K3 can continuously complete long-term engineering tasks, understand and process large code bases, and coordinate the use of terminal tools with minimal human supervision. The third is multi-modal understanding capabilities. K3 natively supports multi-modal input of images and text. It is not a plug-in visual encoder, but a true native visual understanding. This means it can handle text, image and video input simultaneously. There are four ways to use the API. Chat directly and for free through the Kimi mobile app (iOS/Android/HarmonyOS). The desktop version of Kimi Work 3.1.0 and above provides work scene support. Kimi Code terminal tool is aimed at programmers. The Kimi API platform is aimed at developers and enterprise users. The official also provides the Kimi Enterprise enterprise version, which provides data privacy protection and member management functions. From the perspective of actual user experience, front-end developers and programmers generally give high evaluations and believe that K3 "understands the design intent accurately" in front-end code generation and UI restoration. However, the feedback from ordinary users is polarized - some people feel that it is indeed better at analyzing complex long texts, while others complain that "there is no huge difference in daily conversations."

  • Kimi K3’s API pricing is at the high end of China’s head models – output is US$15 per million tokens, which is approximately RMB 100. The previous generation K2.6 is $4, which has increased nearly 3 times. The output price of the similar domestic model GLM-5.2 is US$4.4. However, an important data given by the official is: in programming scenarios, the cache hit rate can exceed 90%, so the actual expenditure is far lower than the listed price. Cache hit input costs only $0.3 per million tokens. According to officials, this system that relies on Mooncake’s separated reasoning architecture makes K3 highly cost-effective in high-frequency programming scenarios. Huang Zhenxin, head of Dark Side of the Moon Enterprise Business, made it clear that open source models and large Chinese models should not be labeled as low-priced. "We have built a SOTA-level model that can also match reasonable commercial pricing." Artificial Analysis' calculations show that the single-task cost of Kimi K3 is about US$0.94, which is close to the US$1.04 of GPT-5.6 Sol. It is not much cheaper than it seems on the surface. On the consumer side, the free version of Kimi includes a certain K3 quota, and paid membership is divided into multiple tiers, up to 699 yuan/month. Some users reported that 20% of their quota was used in one day, which means that the cost for heavy users cannot be ignored. Dark Side of the Moon says it will stick to the open source route. Full model weights released before July 27, 2026, under a modified MIT license. For enterprise customers with private deployment and fine-tuning needs, open source is a high-value option.

  • Positive comments mainly focus on three aspects. First, the experience in programming scenarios is excellent, the front-end code generation is of high quality, and the design intent is accurately understood. Vercel founder Guillermo Rauch has also conducted actual testing and verification. Second, its ability to process long texts and long tasks is outstanding. Users reported that when putting dozens of pages of documents into a summary, K3 can accurately extract key information, proactively subtract, and "begin to have better judgment." Third, the automation capabilities in professional scenarios (financial analysis, contract review, industry research) are satisfactory. Some users have used it to find hidden automatic renewal clauses in contracts. Negative reviews are also evident. The most common complaint is the slow response speed - a similar task, Claude Fable 5, took 3.5 minutes, while K3 took 12 minutes. API responsiveness is below category average. Secondly, the pricing is on the high side. From K2.6 to K3, the output price has increased nearly three times. Again, there is the problem of "over-reasoning". Some users report that K3 will expand the scope of tasks on its own initiative with vague intentions, increasing rework costs. There is also user feedback that no significant improvement can be felt in daily light scenarios.

  • The industry media generally spoke highly of it. The Economic Daily interpreted this as a new node for China's large models to move toward pricing power. A Goldman Sachs research report believes that China's large models have previously entered the global market mainly based on cost performance, and in the future they may rely on cutting-edge capabilities to enter the core workflow of global developers. Morgan Stanley is concerned that the API price of Kimi K3 is at a high level among China’s top models, and believes that higher pricing has positive significance for the market structure of large models. The tech community is divided. Some developers believe it is "a turning point for the open source model in the field of programming" because it is the first time that an open source model has completely surpassed all closed source models in the front-end code competition. Others believe that Benchmark belongs to Benchmark, and the "slowness and expensiveness" in actual use are unavoidable objective costs. Some industry experts also pointed out that the real value of K3 is not only experienced on the C-side, but that it will serve as a base model to empower a large number of small and medium-sized enterprises. After the model weights are open sourced, a large number of domestic companies do not need to train large models from scratch. In terms of capital markets, U.S. AI stocks saw a wave of declines after the release of K3, reflecting investor concerns that open-source cutting-edge models may impact closed-source pricing capabilities. Elon Musk left two comments about "impressive" in related comments, and listed K3 as a key target for Grok 4.6.

  • The distillation charge was the main controversy. The White House directly made the accusation, but technical experts generally believe that this accusation is untenable. American AI technology expert Nathan Lambert publicly wrote that people who believe that Chinese AI laboratories can produce excellent models simply by stealing intellectual property "it's time to wake up." OpenAI head of strategic future Dean Ball also admitted in the controversial post that K3 performance cannot be achieved by distillation. In terms of technical risks, the slow inference speed of the model is currently the most prominent shortcoming. For production environments that require fast response times, K3 may not be the best choice. In addition, the model has a tendency to "over-reason" and is prone to proactively expand the scope of tasks without explicit instructions, which poses risks in scenarios that require strict boundary control. There are real issues around compliance and data security. The data in Kimi's hosting service is processed within China, which is an important consideration for overseas users with data export restrictions. Although the open source weights can be deployed locally after release, the deployment threshold of the 2.8 trillion parameter model is extremely high - the official recommendation is a super node configuration of more than 64 accelerators, which is basically unaffordable for ordinary enterprises and individuals. The problem of model illusion has not been completely solved either. Huang Zhenxin himself admitted that "hallucinations cannot be completely eradicated." However, K3 has significantly reduced its incidence and requires the fact-checker sub-Agent in the Agent Swarm architecture for secondary verification.

  • People who are suitable for using Kimi K3 are: front-end developers and full-stack engineers, especially teams engaged in complex UI development and front-end engineering; AI Agent engineers and researchers who need long-term automated programming; financial analysts, industry researchers and scientific researchers, who deal with very large documents and conduct long-term in-depth research scenarios; enterprise customers with privatized deployment requirements can perform local fine-tuning and deployment after the weights are open sourced. People who are not suitable for immediate use are: ordinary users who only need daily simple Q&A and copywriting (many affordable models are sufficient); production environment users who require extremely high response speed; individual developers who do not have enough budget or GPU resources, and the threshold for self-deployment is too high. In terms of usage recommendations, you can adopt a layered strategy: use K3 for the 15%-20% of complex tasks (long-cycle programming, multi-step agents, in-depth analysis of long documents), and use lighter models such as K2.7 Code or DeepSeek V4 Pro for daily high-frequency simple tasks. Optimizing API calls based on caching strategies can also significantly reduce costs.

  • Kimi K3 is a landmark product. It allows the label "open source model" to truly compete with top closed source models for the first time. Its breakthroughs in programming and long-range tasks are real, and it has also caused huge repercussions at the industry and market levels. But overall it is still a "scientific player" - it performs top-notch in specific scenarios, but there is still a gap in general scenarios. Its real long-term influence may not lie in the number of API calls today, but in the fact that after the weight is officially open sourced in a few days, it will serve as a base model to empower the entire AI ecosystem.

User Reviews

  • avatar
    CarolynBakerJr
    K3 的前端代码能力是真的强,丢了个复杂的企业后台需求进去,生成出来的页面审美在线,交互逻辑也基本对的。不过速度是真的慢,等得有点着急。

  • avatar
    goldenlion904
    用 K3 做了个 3D 游戏 demo,效果超出预期。它自己从零写代码、修 bug、调样式,全程不用我怎么管就是每次等它改完都要喝杯咖啡。

  • avatar
    Brjen
    API 价格从 K2.6 的 4 刀涨到 15 刀,涨幅快 3 倍。虽然缓存命中能省不少,但心里还是有点难受。

  • avatar
    程彤云
    实测了三天,K3 最大的变化是开始有判断力了。以前问个问题给你列七八条,现在直接砍掉那些正确的废话,留下的全是真正决策有关的信息。这点比参数提升更难得。

  • avatar
    KOgar
    得承认这是国产开源模型第一次跟 Fable 5 和 GPT-5.6 Sol 站到一个台面上。以前说对标只是口号,这次评测数据是真的接近了。

  • avatar
    SandraHart168
    同一个 bug 修复任务,Claude Fable 5 用了 3.5 分钟,K3 用了 12 分钟。差距很明显,不光是速度,对话响应也偏慢。

  • avatar
    游客_8
    办公场景测了写周报和做 PPT。写周报真的香,把零散的工作记录丢进去,自动给你包装成专业版本。做 PPT 就比较弱了,只给文字大纲,页数还失控,不如豆包的一键生成好用。

  • avatar
    Madison_Hall_7
    精度确实有提升,中文综合从 72.9% 到 75.6%,coding 从 62.6% 到 70.3%。但 token 消耗也上去了,同样任务花的钱多了不少。只能说 SOTA 级模型配 SOTA 级价格吧。

  • avatar
    Karen836
    K3 有点太爱表现了,让它总结一段产品说明,它不光总结,还自作主张加了一段风险分析。问题是原文根本没提风险,纯属脑补。复杂任务用它,简单任务还是切回 K2 更省心。

  • avatar
    月光_21
    最搞笑的是官方自己都承认综合方面还落后 Fable 5 和 Sol,但社区吹得好像已经全面碾压了。理性看待吧,K3 在特定场景下确实强,通用能力还是有差距。

  • avatar
    redzebra695
    前端代码是真的强,Frontend Code Arena 1679 分登顶不是吹的。跟 Fable 5 和 Sol 同时生成同一个 SaaS 落地页,K3 的审美确实在线,按钮间距和配色比例看着更舒服。

  • avatar
    AbigailMitchell_2024
    把几十页的行业白皮书扔进去,让 K3 提炼三个关键变化,它给得很准,而且主动标出了需要留意风险的地方。长文档处理确实比 K2 强一个档次。

  • avatar
    realSanjaSilić_dev
    用 K3 检查租房合同,帮找出一条自动续约的隐藏条款。这功能说大不大,但省了我一笔可能的损失。

  • avatar
    VEcoo
    这个模型在科研编程的场景太适合了。我用来复现一篇计算化学的论文结果,K3 自己读了代码库、修了依赖冲突、跑了参数扫描,整个过程跑了大概 5 个小时,但我基本没干预。

  • avatar
    HashHunter114
    API 接入挺简单的,兼容 OpenAI SDK,改个 base_url 就行。速度偏慢但稳定性还不错,跑了几个长任务没有出现中断。

  • avatar
    Olivia.Brooks007
    刚发布就断售了,负载扛不住可见热度多夸张。好在 API 还能用,网页端等了两天才上去。希望扩容快点跟上。

  • avatar
    唐兰
    2.8万亿参数听起来很吓人,但实际激活的只有 16/896 个专家,推理算力要求其实没想象中那么离谱。问题是自部署需要 64 个以上加速器,个人开发者基本没戏。

  • avatar
    Justin.Morales_99
    蒸馏指控真的有点离谱。K3 公布的架构创新——KDA 注意力机制和 Attention Residuals——都是有技术论文支撑的。说靠蒸馏做出 2.8T 参数的模型,技术上就不可能。

  • avatar
    DrErnestoMárquez_x
    编程能力确实强,但推理速度偏慢。同样是生成一个 Three.js 的瀑布场景,GPT-5.6 Sol 几分钟搞定,K3 花了半小时。不过最后成片的质量 K3 反而更好,彩虹和水雾的细节处理更自然。慢工出细活吧。

  • avatar
    blPAR
    有些人的期望太高了,觉得发布后就应该吊打所有闭源。实际用下来,K3 是顶级编程选手但日常对话有过度推理的问题,选对场景很重要。

  • avatar
    Judy_Morales52014
    开通了 699 元/月的最高档会员,重度用了一天额度掉了 20%。虽然心疼但确实值,复杂文档处理效率至少翻倍。

  • avatar
    Richard.RobertsX
    用 K3 做了个流量回放平台,从需求到跑起来大概半天。之前用别的模型至少折腾两天。它对长代码库的理解力确实比前代强很多。

  • avatar
    胡勇丹
    做 2D 游戏很强,让 K3 复刻泡泡堂,它自己拆成两条并行工作线,地图、角色、逻辑一起写,最终跑起来能玩。就一句话需求,它理解得太好了。

  • avatar
    BCooper_2023
    普通用户可能感受不到 K3 的强大,因为日常写文案聊天跟 K2 差别不大。但如果你做编程或者研究,差距是非常明显的。

  • avatar
    DanielleRamirez_668
    马斯克又点赞了,还说要拿 Grok 4.6 跟 K3 对标。能被这个级别的玩家当成对标目标,本身就是实力认证了。

  • avatar
    VincentPerezZ
    刚上手就给它丢了一个大项目,自己写了个迷你 GPU 编译器,从零开始,最后跑通了 nanoGPT 训练。虽然不是生产级的,但 48 小时能做成这样确实让人惊讶。

  • avatar
    吴秀龙
    月之暗面的开放策略值得肯定,K3 权重将在 7 月 27 日前开源。国内中小企业可以基于这个基座做二次开发,不用从零训练大模型了。

  • avatar
    MidhaelJensen
    第一时间就试了,拿前端代码竞技场的题让它跑,确实一把过。这是第一次有开源模型在这个榜单上超过 Claude 和 GPT。