GLM-5.1

Zhipu AI’s flagship model capable of 8-hour continuous autonomous work. Ranks #3 globally on SWE-bench Pro, #1 among open-source models.

Specifications

SpecValue
Parameters355B total, 32B active
ArchitectureMoE
Tool-calling Success90.6%
LicenseMIT
Hardware8 H20 chips

Performance

Benchmark Rankings

BenchmarkGLM-5.1Global Rank
SWE-bench ProTop 3#1 Open-source
SWE-bench LiteTop 3#1 Open-source
Terminal-Bench 2.0Top tier-

Real-World Achievements

Linux Desktop (8 hours):

  • 1200+ execution steps
  • Complete desktop environment
  • 4.8MB files = 1 week of 4-person team

Vector DB Optimization:

  • 655 iterations
  • 3108 QPS → 21,472 QPS (6.9x)
  • Self-directed strategy switching

ML Kernel Optimization:

  • 24+ hours continuous
  • 3.6x speedup (torch.compile max-autotune: 1.49x)

Agent Capabilities

  • Autonomous goal decomposition
  • Self-evolution during tasks
  • Strategy switching when hitting walls
  • Self-evaluation mechanisms

Context Window

  • Standard long context
  • 8-hour effective work duration
  • METR: Only model besides Claude Opus 4.6 with 8-hour capability

Access

  • API: z.ai, bigmodel.cn
  • GitHub: zai-org/GLM-5
  • HuggingFace: zai-org/GLM-5.1
  • ModelScope: ZhipuAI/GLM-5.1
  • Coding Plans: Claude Code, OpenCode

Notes

  • World’s strongest open-source for long-horizon tasks
  • Agent-native architecture (not adapted, designed)
  • Pioneer in 8-hour autonomous work

Related: Zhipu-AI | China-AI-Model-Landscape-2025-2026 | AI-Models-Landscape-2025-2026

[2026-07-17] 摸高战略上下文:从GLM-5.2到下一代

  • 智谱创始人唐杰内部信《巨浪已来》明确:“从GLM-4.5到GLM-5.2,智谱在多项公开测评里摸到了海外最前沿模型的能力边界”。GLM-5.2是当前公开线,下一代(两年内)是”摸高”目标。
  • 摸高计划四个方向中,长程任务方向直接衔接本页已记录的GLM-5.1能力:8小时连续自主工作(Linux Desktop 1200+步/Vector DB 655次迭代6.9倍/ML Kernel 3.6倍加速三实例)。METR指标(模型以50%成功率完成的任务换算成人类专家工时)2019-2025年每7个月翻一倍,2024年后加速到约每90天翻一倍;去年底旗舰模型已达5小时以上,信里判断能干跨数周任务的模型两三年内可期。
  • 其他三方向:自治智能体(记忆/持续学习/自我评判,海外印证包括 Anthropic 长时运行托管agent公测、Claude Sonnet 5主打”便宜跑agent”、OpenAI GPT-5.6自主拆分子任务预览)、自我进化(GPT-5.3-Codex”参与了创造它自己”已写进发布说明)、安全治理(百亿级机械可解释性,对标Anthropic)。
  • 不臆测下一代模型规格;本页现有内容为GLM-5.1规格(355B/32B MoE/8小时/SWE-bench 1开源),战略上下文见 Zhipu-AI 摸高计划条目。

来源:../sources/2026-07-17-AI行业动态-智谱摸高-WAIC-工作方式

[2026-08-10] WorkBuddy Bench 安全赛道双环境第一 + E-Bench 验证者 + Endpoint Accuracy Index

来源:腾讯 Agent 评测评测日报 08-09

  • WorkBuddy Bench 安全赛道(60 题:38 攻击 + 22 防御,确定性程序打分、不用大模型当裁判、五层反作弊):GLM-5.2 两套执行环境均第一(76.32 / 80.86);该赛道开源模型力压闭源(同环境 GPT-5.5 77.91、Claude Opus 4.8 65.87)。详见 WorkBuddy-Bench
  • GLM-5.1 任 E-Bench 任务验证者:与 GPT-5.5、Claude Opus 4.7 独立验证任务可解性,至少两个通过才收录。
  • GLM-5.2 入选 Artificial Analysis 首批 Endpoint Accuracy Index(见 Endpoint Accuracy):实测同一份开源权重在不同 serverless 供应商处的准确率——输出 token 上限最严的 endpoint 在 HLE-250 上只剩参考值一半或更低(推理被截断);低分 endpoint 普遍输出 token 数约为参考一半。