行业观察

智谱上线GLM-5.3-FlashX:10万张国产卡跑通200 tokens/s高速推理

Published 2026-09-20 Author Source 上海证券报 · https://www.cnstock.com/commonDetail/792135 · 2026-09-18 Tags AI-Native Organization / Self as Product / Launch
新模型推理速度最高达每秒200 Token,定价为原版的2.5倍,API与体验中心同步开放。

智谱9月18日正式推出GLM-5.3-FlashX模型,API与体验中心同步开放。新版本最高输出速度可达每秒200个Token,较现有GLM-5.3-Flash提升约5倍,定价相应调整为原版本的2.5倍。官方表示,新版本在智能、价格、速度三个维度形成了综合竞争力。

基础设施方面,智谱披露GLM-5.3-Flash在两周内完成超过10万颗国产AI加速器集群的部署,吞吐提升3.2倍。同时,智谱发布了公司首个RSI(递归自我改进)成果,实现“10万国产卡用GLM造GLM”的闭环验证。此前GLM-5.3-Flash曾以“Ox Alpha”之名面向全球开发者开放,调用量持续攀升。

【AICOR点评】把“国产算力可训练”从单次验证推进到可持续工程循环,这是国产AI基础设施的关键一步。推理速度的提升不只是用户体验问题,更决定了Agent能在多大程度上进入实时交互、高频调用的业务场景。对要做国产替代的企业来说,软硬件协同的稳定性已经是可以验证的事实。

Copyright Notice This article is AI-assisted, rewritten from public reports. Copyright of the information belongs to the original authors and media; content is for industry sharing only and does not constitute investment or business advice.

For copyright concerns, contact AICOR (400-601-8080 / WeChat: aicor-ai); we will handle it promptly upon notification.

← Back to News