行业观察

DeepSeek发布V4.1-Flash:552B参数MoE,推理成本降至行业新低

Published 2026-09-13 Author Source 新浪财经(招商证券研报) · https://stock.finance.sina.com.cn/stock/go.php/vReport_Show/kind/lastest/rptid/842628788149/index.phtml · 2026-09-13 Tags AI-Native Organization / Self as Product / Launch
DeepSeek推出全新架构的V4.1-Flash模型,KV缓存需求大幅压缩,闲时输入成本低至0.02元/百万token。

9月10日,DeepSeek正式发布V4.1-Flash模型,这是其全新模型结构系列中最小尺寸的原生多模态MoE模型。该模型总参数552B,采用全新的Causal-Encoder-Decoder架构,输入激活仅8B、输出激活16B,成本显著低于同尺寸模型。

在缓存优化方面,V4.1-Flash将KV Cache的HBM需求降至上一代的1/4,SSD需求降至1/8,上下文窗口扩展至1M token。定价采用峰谷机制,闲时缓存命中输入仅0.02元/百万token。多项Agentic基准测试成绩超越前代旗舰V4 Pro,V4 Pro将于9月14日下线。

此外,路透社此前报道称DeepSeek已聘请中信证券筹备科创板上市事宜,计划年内启动IPO进程。2026年前7个月公司营收约4.75亿元,约为2025年全年的10倍。

【AICOR点评】DeepSeek用架构创新而非参数堆砌来压缩推理成本,这条路径对Agent场景尤其关键——缓存命中成本越低,多步骤智能体的经济可行性越高。0.02元/百万token的定价将倒逼行业重新评估"模型调用"的单位经济学。对企业AI落地而言,推理成本下降到临界点后,原本因ROI不达标而搁置的AI流程自动化需求将集中释放。

Copyright Notice This article is AI-assisted, rewritten from public reports. Copyright of the information belongs to the original authors and media; content is for industry sharing only and does not constitute investment or business advice.

For copyright concerns, contact AICOR (400-601-8080 / WeChat: aicor-ai); we will handle it promptly upon notification.

← Back to News