资讯

Ox Alpha 现身:智谱开源 GLM-5.3-Flash 原生多模态模型,限时折扣价为 GLM-5.3 的 1/20

IT之家·2026/8/26 14:21:30🔗 原文

📌 概要

<p>IT之家 8 月 26 日消息,智谱今晚上线并开源 GLM-5.3-Flash (320B-A18B),这是 GLM-5 系列的首个原生多模态模型。该模型总参数量为 320B(3200 亿),激活参数仅 18B(180 亿)。</p><p style="text-align: center;"><img class="no-alt-img" src="https://img.ithome.c

<p>IT之家 8 月 26 日消息,智谱今晚上线并开源 GLM-5.3-Flash (320B-A18B),这是 GLM-5 系列的首个原生多模态模型。该模型总参数量为 320B(3200 亿),激活参数仅 18B(180 亿)。</p><p style="text-align: center;"><img class="no-alt-img" src="https://img.ithome.com/newsuploadfiles/2026/8/441ac132-3ddd-40cc-8581-83f9d385c085.jpg?x-bce-process=image/format,f_auto" /></p><p>智谱在正式发布前,以匿名模型 Ox-Alpha(中文社区称作牛来)在 OpenCode 和 OpenRouter 上进行了大规模测试。Ox-Alpha 迅速成为当周最受欢迎的模型,创下双平台调用量新高。<strong>而这些请求流量全部由国产芯片提供算力支持</strong>。</p><p style="text-align: center;"><img class="no-alt-img" src="https://img.ithome.com/newsuploadfiles/2026/8/288f3357-162f-4679-ab1b-636e68b0f583.jpg?x-bce-process=image/format,f_auto" /></p><p style="text-align: center;"><img class="no-alt-img" src="https://img.ithome.com/newsuploadfiles/2026/8/f06cb86a-f9ea-4f79-b5d1-7a99f422998f.jpg?x-bce-process=image/format,f_auto" /></p><p>GLM-5.3-Flash 能力超过 GLM-5.2,在全球权威的 Artificial Analysis Intelligence Index(AA 综合智能指数)中取得 57 分,进入全球前沿模型能力区间,与 Anthropic 最受欢迎的模型 Claude Opus 4.8 得分持平。在自研 <span class="link-text-start-with-http">Z.ai</span> Code Bench 体感评估中,其编程表现与 Claude Opus 4.8 相当。</p><p style="text-align: center;"><img class="no-alt-img" src="https://img.ithome.com/newsuploadfiles/2026/8/4c5f96b3-3d3d-479a-b425-768edb79a238.png?x-bce-process=image/format,f_auto" /></p><p>与此同时,GLM-5.3-Flash 定价为 GLM-5.3 的 1/10,限时折扣内为 GLM-5.3 的 1/20,为 Opus 4.8 的 1/40。</p><p style="text-align: center;"><img class="no-alt-img" src="https://img.ithome.com/newsuploadfiles/2026/8/f65afa69-52a3-4942-a62b-d9b4f2e1ac49.png?x-bce-process=image/format,f_auto" /></p><p>官方表示,GLM-5.3-Flash 在各项基准测试和实际使用中,全面超越参数量高出一倍的 GLM-5.2,定价却仅为后者的 1/10。</p><p>据介绍,这得益于其全新模型架构。GLM-5.3-Flash 的架构专为极低成本而设计,总参数量与 GLM-4.5 相当(355B vs. 320B),但激活参数量(32B → 18B)与层数(92 → 45)几乎减半。结合最新的 30T Token 多模态预训练语料,GLM-5.3-Flash 能够以更少的计算资源实现更强的性能。</p><p>GLM-5.3-Flash 同时进行了多项架构升级,是首个采用稀疏注意力与线性注意力混合架构的开源前沿模型,在保持精准长上下文能力的同时,大幅降低长上下文服务成本;并采用流形约束超连接(Manifold-Constrained Hyper-Connections,mHC)提升模型 Scaling 能力。</p><p>官方表示,GLM-5.3-Flash 视觉能力被原生融入 Coding 循环,使模型能够主动观察界面、渲染结果与交互反馈,并据此持续测试和改进。无论是前端开发、游戏构建、Blender 3D 场景,还是 BUA、CUA 驱动的真实环境操作,模型都能在代码、浏览器和图形界面之间协同完成任务。</p><p>GLM-5.3-Flash 进一步拓展至 Office、金融研究和专业文档等工作场景。它能够自主拆解复杂目标、调用合适工具、检查并优化输出,完成从研究分析、模型构建到 PPTX、PDF、DOCX、XLSX 成品交付的完整工作流。</p><p>GLM-5.3-Flash 正式面向全球开源,已接入 ZCode 等编码平台,并纳入 GLM Coding Plan(每天限量发放 10,000 张体验卡),同步开放 API 调用。</p><p>IT之家附相关链接:</p><ul class=" list-paddingleft-2"><li><p style="text-align: left;">BigModel 开放平台:<a href="https://docs.bigmodel.cn/cn/guide/models/vlm/glm-5.3-flash" target="_blank"><span class="link-text-start-with-http">https://docs.bigmodel.cn/cn/guide/models/vlm/glm-5.3-flash</span></a></p></li><li><p style="text-align: left;"><span class="link-text-start-with-http">Z.ai</span>:<a href="https://docs.z.ai/guides/vlm/glm-5.3-flash" target="_blank"><span class="link-text-start-with-http">https://docs.z.ai/guides/vlm/glm-5.3-flash</span></a></p></li><li><p style="text-align: left;">GLM Coding Plan:<a href="https://bigmodel.cn/glm-coding" target="_blank"><span class="link-text-start-with-http">https://bigmodel.cn/glm-coding</span></a></p></li><li><p>HuggingFace:<a href="https://huggingface.co/zai-org/GLM-5.3-Flash" target="_blank"><span class="link-text-start-with-http">https://huggingface.co/zai-org/GLM-5.3-Flash</span></a></p></li></ul>