Reka Core 2
Reka AI 第二代旗舰原生多模态模型,支持文本、图像、音频、视频理解,主打企业级可靠交付与 API 易用性
In-depth Report
-
Reka Core 2 is the second-generation flagship multi-modal model launched by Reka AI in March 2026. It natively supports four types of input: text, image, audio, and video, and focuses on enterprise-level reliable delivery and API ease of use. It continues Reka's technical route of training from scratch and native multi-modality, and still retains its differentiated advantages over head-closed source models in video understanding. However, there is still a significant gap in community volume and public benchmark disclosure compared to OpenAI, Anthropic, and Google.
-
Reka AI was founded in 2022 by a group of research scientists from Google DeepMind, Meta FAIR and Baidu, and is headquartered in Sunnyvale, California, USA. The core team includes CEO Dani Yogatama, CTO Cyprien de Masson d'Autume, and co-founders Qi Liu, Mikel Artetxe, etc.; the original chief scientist Yi Tay has left at the end of 2024 and returned to Google DeepMind. The company has adhered to "efficiency" as its core line since its inception, emphasizing the training and deployment of market-leading models at lower computing power costs. In terms of financing, Reka will complete a Series A round of approximately US$58 million in 2023 (led by DST Global), and received another US$110 million in Series B financing in July 2025. Its valuation exceeded US$1 billion and became a unicorn. Investors include Nvidia and Snowflake. The new funds will be used to accelerate multi-modal platform research and development and enterprise-level market expansion. Reka's product matrix is divided into four levels according to the number of parameters: Spark (1B), Edge (7B), Flash (21B) and Core (flagship). Flash 3 will be open sourced under the Apache 2.0 protocol in March 2025.
-
Reka Core 2 is an iterative version of Reka Core, which is officially positioned as the “second-generation flagship multi-modal model.” It has improved on text, image, video, and audio understanding tasks compared to the first generation. Officials say it can compete with front-line models, while focusing on enterprise reliability and API accessibility. The model context window is 128K, covering five types of capabilities: text generation, vision, audio understanding, video understanding and reasoning. It is provided externally through Reka API and AWS Bedrock. Reka's core technical narrative is "native multimodality": text, images, videos, and audio are first-class citizens from the training stage, rather than adapters mounted afterwards. The biggest highlight of this route is video understanding - the first generation Reka Core scored 59.3 points on the Perception-Test video question and answer benchmark, exceeding Gemini Ultra's 58.7 points. At the time, GPT-4o did not yet support native processing of continuous video clips. Core 2 follows this capability orientation and strengthens its integration with agent workflow, which can be used in scenarios such as multi-modal RAG, video retrieval, and event detection.
-
According to Reka's official pricing document (verified in June 2026), the text call price of Reka Core is US$2.00 per million input tokens and US$6.00 per million output tokens; images are charged on a per-image basis (approximately US$0.02/image), video is US$0.08/minute, and audio is US$0.02/minute. For comparison, Edge is $0.10/$0.10 and Flash is $0.80/$2.00. Reka Research agents are billed per thousand requests at $25 for standard, $35 for parallel thinking low/high and $60 for high-end. Compared with the US$10/25 price tag when the original Core was released, the unit price of Core 2 has dropped significantly, reflecting the "efficiency first" pricing orientation. Enterprises can also choose managed cloud, local, VPC or completely offline (air-gapped) deployment, with deployment flexibility close to that of traditional database vendors.
-
Reka’s user reputation is concentrated in two groups. First, there are enterprise customers who require native video/image understanding, such as Shutterstock, Turing Video, etc., who use Reka Vision for large-scale video and image retrieval; second, there are customers in finance, medical and other industries who value deployment controllability. Offline deployment and keeping data out of the domain are core selling points. Developers generally believe that Reka's multi-modal native capabilities are real and usable, and the API and Python/JavaScript SDK experience is smooth. Negative feedback mainly comes from the lack of voice in the community: there are few open benchmarks, third-party reproductions and real use cases, and developers lack enough neutral reference when making selections. In addition, the specific parameter quantities, training details and independent benchmark scores of Core 2 are not disclosed, and the information transparency is weaker than that of competing products, which increases the evaluation cost of the company.
-
The industry generally regards Reka as a "small but sophisticated" multi-modal model manufacturer, recognizing that its small team of senior research scientists can produce competitive models with far lower computing power than giants. The third-party model index (aaas.blog) gives Reka Core 2 high quality and freshness ratings (Quality A, Freshness A+), but low adoption (C), citations, and participation (both F). The comprehensive index is only 33. It is recommended to use it with caution. This is consistent with Reka’s corporate positioning: the model capability is not weak, but the ecological and community maturity has not yet been established. Compared with the head model, Reka Core 2 is close to the same level as GPT-4, Claude, and Gemini in terms of text/image benchmarks (the original Core's MMLU 83.2, GSM8K 92.2, and HumanEval 76.8 can be used as a reference). Video understanding is its relative advantage; its shortcomings lie in the scale of the ecosystem, the richness of tool calls, and external integration.
-
The main risks focus on three points. First, there is a lack of transparency: Core 2’s parameter volume, training data composition, and independent benchmarks are not disclosed, and companies need to conduct sufficient PoC on their own to assess real capabilities. Second, fluctuations in talent and ecology: Chief Scientist Yi Tay will leave at the end of 2024, and the stability of the core team is a long-term variable. Third, the low market volume results in a scarcity of community resources, troubleshooting documents, and Chinese materials, and the cost of getting started is high for individual developers or small teams.
-
Reka Core 2 is suitable for enterprise customers who require native multi-modality (especially video/image understanding) and have strict requirements on data residency and deployment methods, such as media content retrieval, security event detection, multi-modal customer service and document analysis. It's also suitable for teams that are already using AWS Bedrock and want to add multimodal capabilities without having to onboard a new vendor. It is not suitable for individual developers who pursue the largest community ecosystem, the richest tool chain, or require a large number of public benchmark endorsements; if the scenario is mainly pure text reasoning, the ecological maturity of the same closed-source model is usually higher. Alternatives include the GPT-5.6 series, Claude Opus, Gemini series, and open source multi-modal models such as Qwen-VL and InternVL.
-
Reka Core 2 is an enterprise-level native multi-modal flagship model with solid capabilities and flexible deployment. Video understanding is its differentiation, but ecological scale and transparency are still lessons it must make up for before it can move into a more mainstream market.
User Reviews
-
Noah.MooreIII—整体是能力扎实的企业级多模态模型,视频是长板,想要最全生态还是去抱 OpenAI。 -
Stephen_RamirezQ—说实话 Core 2 具体参数和独立基准都不公开,选型时只能自己跑 PoC,透明度这块确实扣分。我们内部对比了一圈,多模态尤其视频确实强,但纯文本推理离一线旗舰有差距。如果你的核心需求就是视频和图像理解,它值得上,否则建议先看清楚再决定。 -
JoshuaAllen_2021—没有 ChatGPT 那种聊天界面,纯 API,想给团队用得自己搭前端,门槛摆在那。 -
HaroldFoster77—我们安防场景用它做摄像头事件检测,连续视频理解比抽帧方案误报少很多。之前用别的模型得先把视频切成一帧帧再识别,丢了大量时序信息,Core 2 原生看视频之后,越界、逗留这类依赖前后文的行为识别准了不少,落地效果超预期。 -
BrianKelly_99—Reka Core 2 的视频理解是真能打。我们做体育集锦自动剪辑,之前用抽帧方案经常漏掉连续动作,换成 Core 2 直接喂整场录像让它按时间戳出高光片段,命中率高了一大截,后期人工补的工作量明显少了。 -
xAmáliaCastro_x—首席科学家 Yi Tay 都回 DeepMind 了,长期团队稳定性是个问号。 -
ATurner—在物理 AI 和机器人场景里 Reka 是真首选。它不像很多大模型把视频当一堆静止图片处理,而是真的在统一潜空间里理解动作序列和意图,时序推理明显更强。我们做机械臂演示,让它看一段操作视频就能复述步骤,这种能力别家还真不一定有。 -
Brittany.Carter00761—价格还行,Core 2/6 每百万 token,比 GPT 同档便宜一截,但视频按分钟另算,量大要算清账单。 -
Lauren_PatelII6—n8n 社区节点出来了,媒体团队用 Reka Clips 工作区做短视频切片方便不少。 -
Deborah_Sanders_752—视频问答偶尔会一本正经胡说八道,关键场景还是得人工复核,不能完全托管。 -
AngelaReed_202199—企业部署灵活度是它最大的卖点。我们金融客户的数据必须不出域,Core 2 支持本地和完全离线部署,合规那边一次就过了,这点比 OpenAI 和 Claude 好谈太多。代价是得自己维护推理栈,小团队没运维的话还是走云 API 省事。 -
Sharon_RuizZ09—通过 AWS Bedrock 也能调,已在用 Bedrock 的团队接多模态不用再接新供应商,省事。 -
GGonzalezJr—代码能力还行,但 prompt 得写仔细,它特别较真字面意思,新手容易翻车。 -
WRivera_202270—128K 上下文在 2026 年属实有点寒酸,Google 都给到 2M 了,长文档 RAG 只能切片。 -
Alice_Fisher8—Reka Research 并行思考档偏贵,最高 60 刀每千次请求,值不值见仁见智。 -
DonnaRoberts—原生多模态不是吹的,图像音频视频文本统一处理,内容审核一条管线通吃。 -
Alan.Thompson_Max—二十来人的小团队做出这水平不容易,效率路线值得respect。 -
RachelBellQ—生态太小了,中文资料基本搜不到,排障全靠翻文档,个人开发者慎入。