Z.ai released GLM 4.6 on September 29, 2025. The registry currently records it as open weights.
Compared with GLM-4.5, this generation brings several key improvements:
Longer context window: The context window has been expanded from 128K to 200K tokens, enabling the model to handle more complex agentic tasks. Superior coding performance: The model achieves higher scores on code benchmarks and demonstrates better real-world performance in applications such as Claude Code、Cline、Roo Code and Kilo Code, including improvements in generating visually polished front-end pages. Advanced reasoning: GLM-4.6 shows a clear improvement in reasoning performance and supports tool use during inference, leading to stronger overall capability. More capable agents: GLM-4.6 exhibits stronger performance in tool using and search-based agents, and integrates more effectively within agent frameworks. Refined writing: Better aligns with human preferences in style and readability, and performs more naturally in role-playing scenarios.
The structured profile below stays connected to the canonical model record, including its architecture, parameters, papers, benchmarks, media, licence, and source links as those fields are enriched.
Model publication
zai-org
GLM 4.6Open the source for GLM 4.6
Compared with GLM-4.5, this generation brings several key improvements: Longer context window: The context window has been expanded from 128K to 200K tokens, enabling the model to handle more complex agentic tasks. Superior coding performance: The model achieves higher scores on code benchmarks and demonstrates better real-world performance in applications such as Claude Code、Cline、Roo Code and Kilo Code, including improvements in generating visually polished front-end pages. Advanced reasoning: GLM-4.6 shows a clear improvement in reasoning performance and supports tool use during inference, leading to stronger overall capability. More capable agents: GLM-4.6 exhibits stronger performance in tool using and search-based agents, and integrates more effectively within agent frameworks. Refined writing: Better aligns with human preferences in style and readability, and performs more naturally in role-playing scenarios.
- Licence
- mit
- Architecture
- Glm4MoeForCausalLM
- Parameters
- 357B
- Context
- 204.8K tokens
- Artifact format
- safetensors
- Quantization
- native BF16
- Download size
- 665 GB

Hardware to run it (inference, estimated)
798 GB VRAM minimum · 997 GB recommended · 997 GB system RAM · 665 GB storage
Estimated from parameter count and stored precision; verify against the selected runtime and context length.
Benchmarks
No benchmark observation recorded for this model.