iFlyTek Spark 6
科大讯飞based on 全国产算力训练的认知large model,语音 and 教育两大场景极强,政企落地领先,但通用对话 and 代码能力偏中下
In-depth Report
-
"iFlytek Spark 6" is an entry included in the product express, but one thing needs to be made clear first: as of July 18, 2026, iFlytek has not officially released a version named "Spark 6" or "Spark V6". The version pedigree that can be verified on the official website and Baidu Encyclopedia is V1.0→V1.5→V2.0→V3.0→V3.5→V4.0→V4.0 Turbo, and then moved to the deep inference large model "Spark X1" and its upgraded version, as well as the native multi-modal large model "Spark X2-VL" released in June 2026. In other words, there is no generation "6" in the official naming system. This report is written based on the verifiable current main models (Spark X1 upgraded version + X2-VL multi-mode + Spark 4.0 Turbo base). "Spark 6" is more likely to be the expected name of the next generation by the media or channels. The final naming and classification are left to manual review and confirmation in the background. Putting aside the name, iFlytek Spark’s real background lies in three points: it is extremely strong in the two main fields of national computing power training, voice and education, and leads the country in government and enterprise landing orders; its shortcomings are also clear, and its general dialogue, writing creativity and coding ability are below average in the first echelon of domestic products.
-
Founded in 1999, iFlytek is a veteran leader in the field of domestic intelligent speech. In 2008, it received a hearing score close to that of a real person in the International Speech Synthesis Competition. Later, it expanded its technology landscape from speech recognition and speech synthesis to cognitive intelligence and general large-scale models. The Spark large model officially released V1.0 on May 6, 2023, and has maintained high-frequency iterations since then. It is one of the first batch of domestic large models in China to benchmark ChatGPT. The most distinctive strategic label of this company is "full stack independent and controllable". It teamed up with Huawei to build the domestic Wanka computing power platform "Feixing 1" based on the Ascend ecosystem. In 2024, it launched the larger-scale "Feixing 2". In 2026, it will further promote national production computing power training based on Ascend 950. Chairman Liu Qingfeng has a saying that has been quoted repeatedly: "If we were not using a domestic platform, but an already established foreign platform, the effect of today's Spark model might be better. But you have to take this step, unless you don't want to be independent and self-reliant." In the context of the United States continuing to tighten the export of AI chips, "using domestic computing power to complete training" rather than just "running inference on domestic computing power" has become the moat that distinguishes iFlytek from most of its peers. The operational level also supports this narrative: the company's overall revenue in the first half of 2025 was 10.91 billion yuan, and the company won 210 projects related to large model bidding throughout the year, with a total amount of approximately 2.316 billion yuan, ranking first among domestic general large model manufacturers. Most of its customers are government and enterprise entities such as education bureaus, public hospitals, and energy state-owned enterprises.
-
Spark adopts MoE hybrid expert architecture. The official website claims that the reasoning efficiency has been improved by 100%. It focuses on seven core capabilities: text generation, language understanding, knowledge question and answer, logical reasoning, mathematical ability, coding ability, and multi-modal interaction. Around the base, the product line is spread out. On the personal side, there are Spark App, Spark PC version, and SparkDesk desktop version. After the upgrade in September 2025, it will focus on four things: "AI writing, problem solving, AI reading, and in-depth research." In-depth research can open up the link from "searching for information to publishing PPT" at once. Voice is its specialty: multilingual speech recognition covers 37 languages and more than 130 language capabilities, and dialect recognition is almost unmatched in China; "Xiaoxing Chat" supports full voice, visual, digital human multi-mode interaction, responds within 1.5 seconds, can interrupt at any time, and can sense emotions. In terms of multi-modality, X2-VL, released in June 2026, is currently the only mainstream multi-modal model based on national industrial computing power training. It can do everything from looking at pictures, reading documents, analyzing charts, and understanding pictures. In the college entrance examination simulated multi-modal test questions, the accuracy rate in all subjects is close to 95%. On the enterprise side, iFlytek released four new enterprise-level products in June 2026: Spark Marketing Assistant (AI marketing digital employee, supporting "1 real agent + 5 AI clones" collaboration), Spark AI Enterprise Risk Control Officer (risk physical examination, one-click risk report, efficiency increased by 5 to 8 times), Spark Intelligent Voice VoiceWise (high-performance CPU version of ASR without GPU, full-duplex interaction), Spark Intelligent Reasoning Platform SparkInfer (four-level model routing, semantic caching, and token management, combined can save 30% to 60% tokens). The ecological implementation is quite intensive. Huawei Mate70's call summary, BYD's smart sparring, Changhong AI TV, Hikvision's iFlyCode code assistant, and PetroChina's Kunlun large model all use Spark capabilities. The real experience has mixed reviews. The scenes of mathematics, education, and pronunciation are indeed eye-catching - according to actual tests conducted by many media in the 2026 College Entrance Examination, Xinghuo ranked first with a score of 148/150 in mathematics paper I of the new college entrance examination, and ranked first in Shanghai paper composition with 65.5 points. However, general conversation and writing are recognized weaknesses. Many reviews ranked it at the bottom of the first echelon of domestic products, and its coding ability was criticized as "it can make mistakes even in the order of writing."
-
The personal version as a whole follows the "free traffic" route. The App and web version are free to use, and the free version for writing has a quota of about 5,000 words per day. The API side is where iFlytek engages in the most severe price war: In May 2024, it announced that the Spark Lite API is "permanently free", making it the first company in the industry to make the basic version permanently free; the Pro/Max version is as low as 0.21 yuan/10,000 tokens. According to iFlytek's caliber of "1 token≈1.5 Chinese characters", a content of the size of "Alive" can be generated for about 2.1 yuan, and the price is about Baidu Wenxin at the time ERNIE-4.0, one-fifth of Alitonyi Qwen-Max. The business logic is very clear: basic capabilities are free in exchange for traffic and developer scale, and high-value customization, privatized deployment, and industry fine-tuning are charged. This "funnel-style" approach has been verified on the iFlytek open platform, linking more than 8 million developers. What really makes money is the overall solution for government and enterprises - in the first half of 2026, the entire set of industry solutions accounted for more than 60% of iFlytek's large model revenue. The amount of a single contract continues to rise. What is sold is no longer the model itself, but the complete delivery of "scenario + capability + platform".
-
Positive reviews are highly concentrated in two areas. The first is voice. "Speech-to-text is the best." "Mantras can be automatically filtered during meetings, and the accuracy is 30% higher than that of the mobile phone." "Dialect areas can be understood." This is the reputation gained after twenty years of accumulation. The second is education. The specialized education version can correct essays, tutor homework, and generate exercises. Teachers and parents find it practical. The conversation style is also considered to be closer to ordinary people's speech than DeepSeek's "Xueba accent." Negative feedback is equally blunt. Multiple actual tests ranked Xinghuo's general ability at the bottom of the domestic market: "The logic of writing is so messed up that it can lead to grandma's house when writing argumentative essays." "The reasons for recommending 10 books are all copy and paste." "Complex tasks are easy to collapse, and the coding ability is weak." "The UI is too ugly." What is even more worrying is the issue of "sense of presence". Some critical comments directly said, "If I hadn't written this article, I would have almost forgotten that this person existed." "It's good to make PPT, but other times it's just a decoration." Taken together, the consensus among users is that it is the first choice for voice and education tasks, but it is the wrong tool to use for general creation and coding.
-
The evaluation of iFlytek from an industry perspective is much more positive than the C-port reputation, because it looks at different dimensions. The media and analytical institutions generally recognize the certainty of its route of "nationally produced computing power base + general large model capabilities + deep exploration of industry scenarios" - China Daily calls it "the large nationally produced base model that best understands China". The logic is that a highly Chinese, standardized, and evaluable task scenario such as the college entrance examination can test whether the model truly understands the Chinese context and evaluation standards. On the technical side, iFlytek has increased the long-thinking chain reinforcement learning training efficiency from 30% to 84%, and the MoE full-link training efficiency to 93%. These engineering figures point to "domestic computing power moving from being usable to being easy to use." In the competitive product landscape, iFlytek’s position is quite special. In comprehensive capabilities, it is often ranked behind DeepSeek, Tongyi Qianwen, Zhipu, Wenxin, and Doubao; but in the subdivision of "voice interaction", almost all Hengpings rank it as the first choice. This shows that iFlytek is not competing head-on with others on the general list, but relying on differentiated scenarios and independent and controllable narratives such as voice, education, and government and enterprise.
-
The first and most important thing to note is the naming controversy. "Spark 6" is not an officially released version. Currently, the upgraded version of Spark X1 and X2-VL multi-modal are available. During the backend display, manual confirmation is required whether to use the name "Spark 6" or change it to the accurate official model to avoid misleading users. The second is the real shortcomings of general capabilities. Stripped of the propaganda rhetoric, Spark has been repeatedly proven to be weak in the three areas of general dialogue, creative writing, and coding, and the illusion of "serious nonsense" has also been named (although the X1 upgraded version claims to lead the industry in illusion management, there is still a gap in actual experience). If you use it as an all-rounder, you will most likely be disappointed. The third is the gap between "high scores" and "daily". Scores such as 148 points in college entrance examination mathematics and 95% in multi-modality are performance in standardized test scenarios and do not completely correspond to users’ daily casual experience. A strong list does not mean a good feel. This is a question that all models that rely on benchmark narratives have to face. The fourth is C-side presence and product polishing. Many reviewers mentioned that Spark's UI aesthetics, interactive details, and update rhythm are weak. Although "it feels like an abandoned project" is an emotional expression, it reflects that iFlytek's resources are obviously more biased towards government and enterprises rather than the mass consumer market.
-
The people it is suitable for are very clear: users who need high-precision speech transcription, dialect recognition, meeting records, and real-time translation; educators, students, and parents (composition correction, problem-solving tutoring, and oral practice); government and enterprise customers in government affairs, finance, energy, justice, and medical care who have independent controllable, privatized deployment, and data that does not go out of the domain; and developers who value API cost-effectiveness and want to use permanently free Lite or low-priced Pro/Max for lightweight integration. The people who are not suitable are also clear: heavy content creators who pursue top-level general dialogue and in-depth creative writing, code-oriented developers, and ordinary users who want the ultimate C-side interactive experience - for these types of needs, DeepSeek and Tongyi Qianwen are more stable choices in current reputation. The reasonable usage is "scenario-based mashup": leave voice and education to Spark, and leave general Q&A and coding to others.
-
iFlytek Spark is a player that is "strong at home but weak at general use": it relies on its two trump cards of national computing power training, voice education, and intensive government and enterprise implementation to gain a foothold. However, general conversation, writing, and coding, which are the most commonly used abilities by the public, are still at the bottom of the first echelon. The so-called "Spark 6" is currently more of an expectation for the next generation than an official reality. What is really running is the upgraded version of X1 and X2-VL. For government, enterprise and users in specific scenarios, it is a deterministic choice of domestic base; for ordinary C-side users, it is more like a "special tool" than an "all-round assistant".
User Reviews
-
Jonathan_Perry_2021262—语音转文字是真的王炸,开会开着它,老板那些嗯啊的口头禅都能自动过滤掉,准确率比手机自带的高一大截,这块讯飞做了二十年不是白做的。 -
LarryMiller—先说一句,官网压根没有什么星火6,最新的是X1升级版和那个X2-VL多模态,别被名字忽悠了。 -
MarkHendersonSr511—方言识别yyds。 -
Susan_Phillips_7—让它写议论文,论点能给你跑偏到姥姥家,逻辑乱成一团麻,写东西真别指望它。 -
DeFiWav_e—高考数学148分这个是真牛,新京报请了两位特级老师阅卷打的分,六款模型里第一,这推理能力放数学场景确实能打。 -
Frances_PerezSr—作为老师说句公道话,星火的教育版是真好用,批改作文、辅导作业、生成练习题一条龙,家里有娃的可以冲,比通用大模型贴合教学多了。 -
grxxtvrd—免费的谁不爱啊。 -
Brian.Wood—代码能力属实拉胯,让它写个排序都能给我整出bug,搞开发的还是老老实实用DeepSeek吧。 -
浮生_8—UI真的丑,个人审美接受无能。 -
Henry_Evans_X—永久免费的Lite API是真香,做个小项目集成完全够用,Pro/Max也才0.21一万token,比文心通义便宜太多了。 -
FNX816758—我一直觉得讯飞最被低估的点是全国产算力训练,不是拿国产卡跑推理那种,是从头训出来的。中美芯片这么卡的当口,能自己训才是真自主可控,这一步政企客户其实很看重。 -
TerriLowe—怎么说呢,存在感是真的低。 -
heavypanda424—X2-VL那个多模态还行,扔张带图表的试卷进去它能读懂,文档分析也可以,比之前图文分离强多了。 -
MsMilleRasmussen_pro—有没有人知道教育版单独怎么收费啊,看着功能不错但价格好像不便宜。 -
MHughes_995—综合能力横评里它基本垫底,DeepSeek通义智谱文心豆包排完才轮到它,但语音那一栏永远是它第一,这画风也是很割裂了。 -
梁宇勇—做PPT挺好,其他时候基本就是个摆设,我手机里装着但一个月打不开几次。 -
EAbak—口语化这点我给好评,比DeepSeek那种学霸腔听着舒服,更像正常人说话。 -
Christine_Alvarez_2020—实测一圈下来我的结论就是:语音和教育交给星火,通用问答和写代码交给别家,各干各擅长的活最省心,非要拿它当全能选手用的迟早失望,这不是一款让你惊艳的通用模型,但在它的主场几乎无对手。 -
Bobby.HillX—深度研究那个功能试了下,搜资料到出PPT一条龙,摸鱼党狂喜。 -
RonaldDiaz168656—幻觉还是有的,别看官方说X1升级版幻觉治理领先,实际用起来一本正经胡说八道的时候也不少,重要信息还是得自己核。 -
杜燕—讯飞2025年招投标拿了210个项目、23个多亿,国内通用大模型厂商里排第一,客户全是教育局医院能源国企。它根本没打算跟你在C端卷,人家闷声在政企赚钱,这商业模式其实很稳。 -
Sarah_TurnerZ—翻译耳机里用的就是它的技术,出国实测面对面翻译挺顺的,这种软硬一体的场景讯飞是真有积累。 -
DPatelII—医疗V3.5那个版本据说评测榜第一,能多人录音自动生成病历,这种垂类深耕才是讯飞的护城河吧。 -
叶珍英—栓Q,充了会员写作质量也没见涨多少。