LongCat
Meituan’s self-developed 560 billion parameter open source MoE large model platform supports multi-modal capabilities such as text, images, and videos.
In-depth Report
-
LongCat (My Neighbor Totoro) is a self-developed generative AI large model released by Meituan on March 21, 2024. The LongCat-Flash-Chat online experience platform was officially launched in August 2025. The model uses a mixed expert (MoE) architecture with a total parameter of 560 billion. The innovative zero-computation expert mechanism (Zero-NVE) can significantly reduce training and inference costs. LongCat has excellent performance in programming, instruction following and agent tasks, and is positioned as Meituan’s internal AI infrastructure while providing services to external developers through an API open platform. In April 2026, Zhipu and Meituan signed a strategic cooperation agreement to jointly promote AI technology innovation.
-
LongCat is developed by Meituan’s Visual Intelligence Center (VIC), and its core technology is derived from Meituan’s self-developed generative large language model. The LongCat base model will be released in March 2024, and the LongCat-Flash-Chat dialogue platform will be launched in August 2025. Meituan regards AI as a core capability to improve internal operational efficiency, and LongCat supports the intelligent upgrade of Meituan’s multiple business lines such as food delivery, in-store delivery, and hotel and travel services. In addition, Meituan also provides model services to developers through the API open platform. In April 2026, Zhipu AI signed a strategic cooperation agreement with Meituan to cooperate in areas such as complementary model capabilities and API interoperability.
-
The core capabilities of the LongCat platform cover four major directions: text dialogue and creation, which supports daily assistant scenarios such as knowledge Q&A, text generation, writing polish, data analysis, and mathematical calculations; image generation and editing, which supports visual creations such as text-based diagrams and pictures-based diagrams; video generation, which supports the intelligent generation of short video content; and voice calls, which support multi-modal interaction of voice dialogues. LongCat-Flash-Chat also provides deep thinking mode (LongCat-Think) for complex reasoning tasks. The platform's API for developers supports batch calls and tool calls, and can be embedded in third-party applications.
-
LongCat adopts the MoE (Mixed Expert) architecture with a total parameter of approximately 560 billion. The Zero Computing Expert Mechanism (Zero-NVE) reduces reasoning costs and improves computing efficiency by dynamically activating sub-models. The model supports 128K context windows and is able to handle long texts and multi-turn conversations. In terms of programming capabilities, LongCat has surpassed benchmark models such as GPT-4 in multiple authoritative evaluations. In the multi-modal direction, LongCat processes text and images at the same time, supports mixed input and output of graphics and text, and achieves true multi-modal understanding and generation.
-
LongCat adopts an open source strategy, and the model weights have been released as open source on the Hugging Face and GitHub platforms. Developers can obtain it for free and deploy it locally. This strategy has formed a competitive situation with domestic open source models such as GLM and DeepSeek. The commercialization path is realized through API call charging and privatized deployment services, providing customized model fine-tuning and deployment support to enterprise customers. The open source community has given positive reviews to LongCat’s technological innovation.
-
LongCat's fast response speed is the highlight most frequently mentioned by users. In multiple third-party evaluations, LongCat's performance on programming tasks is comparable to GPT-4, showing strong code generation and reasoning capabilities. Meituan internal employees said that LongCat has replaced the original AI auxiliary tools in multiple business scenarios and improved work efficiency. The integration of multi-modal capabilities is also one of the reasons why users praise it. The main pain points are: LongCat is currently mainly for internal enterprises and cooperative developers, and the entrance for ordinary users to directly experience it is relatively limited; the capabilities of the open source version and the online version may be different, and developers need to have certain technical capabilities to deploy and use it.
User Reviews
-
tub0594xm—It’s really nice to have free caching and save a lot of money! -
戴秀—The speed is too slow and I am tired of waiting. -
Debra329—I connected LongCat-2.0 to Claude Code and ran it for a week. The cache hit rate can reach more than 90%, and the cost is indeed low. But it’s still not easy to fix complex bugs. -
Emma_Murray168—I really didn’t expect that the domestic computing power reached trillions of models. The 1M context is so convenient for code analysis. -
6m938ys—After using LongCat-2.0 for two weeks, the overall feeling is that it is suitable for background batch tasks. The strategy of free caching is really smart. The cost advantage is too obvious when repeatedly adjusting tools in OpenClaw. But if you want a chat assistant that responds quickly, forget it. The thinking time only lasts for tens of seconds, and I can’t bear to be impatient. -
Joyce_GomezQ—Owl Alpha deceived the whole world and turned out to be Meituan’s takeout model. It’s hilarious. -
Alice_FosterII—The running score on SWE-bench looks really good, a little higher than GPT-5.5. However, third-party testing said there is a gap between the actual experience and the running scores, so we should pay attention to this. -
DGonzales_7—My coding ability is generally good. I tried it in Codex and asked it to reconstruct an old plug-in. It can analyze the architecture and generate new code by itself, and compile it once. But when I asked it to fix a small bug, it was stuck for a long time. This contrast is quite confusing. It feels like the strategy still needs to be optimized. -
RogerMitchell—Giving away 10 million in free tokens is really generous. -
Pamela.Morales520427—It is a good thing that MIT is open source, but the weight has not been fully released yet, right? It is said to be MIT in theory. Moreover, all benchmarks are vendor-reported, and independent verification has not yet been seen. -
张睿云—Use LongCat to automatically write weekly reports and daily news summaries. It runs in the background without waiting. The cost is almost zero, which is very convenient. -
Brittany_Williams168—After testing LongCat-2.0 for a week, the advantages and disadvantages are very extreme. Advantages: 1M context is enough to cover the entire project, and the free cache makes long-term use costs extremely low. SWE-bench Pro 59.5 can really work. Disadvantages: slow response speed, insufficient stability for complex tasks, relatively dull creative writing, and does not support multi-modal viewing of images. It is recommended to use it as an agent backend, not as the main chat model. -
Janew498—Reasoning is really strong in the re-thinking mode, and AIME can get full marks, but you have to be patient when it finishes thinking. -
Eric_FloresSr—I was pleasantly surprised by its 3D generation capabilities. I can generate interactive Three.js scenes in just one sentence, and the effect is pretty good. It is enough for prototype demonstration. -
菊花_4—The free tokens cannot be used up, and the cache is not charged. It is suitable for raising horses and shrimps. -
阎丹刚—As a developer who writes code every day, I think the greatest value of LongCat-2.0 is that it proves that domestic computing power can be achieved. A 50,000-card cluster can train trillions of models. This engineering capability is really hard. The actual task completion of the agent used to write Python and TypeScript is quite high, but there are often details when writing the front-end UI. If the team is budget-sensitive and needs long-context coding capabilities, it is worth a try. -
JRoss3691—The consumption level is too embarrassing. People who pay 400 yuan a month can’t see it as cheap, but ordinary users can’t afford it if they have a budget of dozens of yuan. The positioning is a bit confusing. -
EthanGonzalez_2022—Open source and free trillions of models, what else do you need a bicycle? -
Amanda848—The cache hit rate is 98%, and using 2.5 million tokens in a day only consumes 180,000. This price/performance ratio is indeed outrageous. It is most suitable to use OpenClaw to run multi-tool cycle tasks. The cost advantage is particularly obvious when tools are repeatedly adjusted in the Agent workflow. But forget about daily conversations, the response speed is impressive, and the impatience is really unbearable. -
KellyColeman_88—To be honest, it is much better than expected. There is a reason why OpenRouter anonymously ranks among the top three in terms of calls in the world. Agent capabilities are close to those of Claude Opus 4.6, and code generation and tool invocation are indeed at the top level in the open source model. However, there are still obvious shortcomings in creative writing. Writing novels is relatively boring, and it does not support multi-modal viewing of pictures, so its application scenarios are limited. -
SawyerWerner—Meituan’s LongCat response speed is really fast, 560 billion parameters are not enough -
STwri—The MoE architecture is indeed efficient, and the zero-computing expert mechanism is very innovative. -
Teresa.Thomas_2021—Programming capabilities can be directly benchmarked against GPT-4, the light of domestic production -
DrEllaWhite—LongCat is open source! 560 billion parameters are fully opened, Meituan is grand -
吴芳—Multi-modal capabilities are well integrated, and both graphics and text can be processed -
4zwain8w—128K context is enough, you can feed long documents as you like -
Bruce_Foster369—LongCat-Think deep thinking mode is very effective in solving complex problems -
KeithBarnesII—The API platform documentation is very clear and easy to access. -
Blake829—Zhipu and Meituan have cooperated, joining forces and looking forward to more capabilities -
BlockWave—Meituan’s internal application of LongCat has achieved good results, and it has also been opened to the outside world.