Compared with GLM-4. 5, this generation brings several key improvements: Longer context window: The context window has been expanded from 128K to 200K tokens, enabling the model to handle more complex agentic tasks.
Model updates refreshed2h agoOct 3, 2026news + changelog
Advanced reasoning: GLM-4.6 shows a clear improvement in reasoning performance and supports tool use during inference, leading to stronger overall capability.
Current rates are not verified. Missing pricing does not mean free usage.
Related provider subscriptions
No current subscription pricing is tracked for this model.
Confirm this specific model, usage limits, and billing terms with the provider. A subscription does not automatically include API credits.
No login needed to compare. Prices are in USD; provider charges are separate from AI Market Cap plans. Context length, caching, tools, taxes, and regional terms can change the final cost. Open weights do not mean free hosting.
41.6
Quality Score
1163
Arena ELO
128.0K
Parameters
205K
Context
Evidence profile
How complete is this record?
This measures the amount of verifiable public evidence we have, not how capable the model is. A missing field means it has not been verified yet, not that its value is zero.
Recent launch, pricing, benchmark, and API signals linked to this model or its provider.
LaunchesZ.ai1mo ago
ZCode now has 1 million users. As a thank-you to our community, we’ve reset usage limits for all GLM Coding Plan users. We’re also rolling out an update that helps turn long-horizon capabilities into
ZCode now has 1 million users. As a thank-you to our community, we’ve reset usage limits for all GLM Coding Plan users. We’re also rolling out an update that helps turn long-horizon capabilities into completed work: - More intelligence in real engineering workflows - 98% cache https://t.co/QT03v4lJVm
RT Cline: One week of @MistralAI's Devstral 2 in Cline. 6.52% diff-edit failure rate. On the diff-edit chart, that puts it ahead of GLM-4.6 (7.58%) an...
Navigation GLM-4.6 Guides API Reference Coding Plan Released Notes Terms and Policy Help Center Get Started Quick Start Overview Pricing Core Parameters SDKs Guide Migrate to GLM-5.3 Models GLM-5.3-Flash/FlashX New GLM-5.3 Hot GLM-5.2 GLM-OCR GLM-ASR-2512 Capabilities Thinking Mode Deep Thinking Streaming Messages Tool Streaming Output Function Calling Context Caching Structured Output Tools Web Search Stream Tool Call Agents GLM Slide/Poster Agent(beta) Translation Agent Vid
Navigation Models GLM-5.3 Guides API Reference Coding Plan Released Notes Terms and Policy Help Center Get Started Quick Start Overview Pricing Core Parameters SDKs Guide Migrate to GLM-5.3 Models GLM-5.3-Flash/FlashX New GLM-5.3 Hot GLM-5.2 GLM-OCR GLM-ASR-2512 Capabilities Thinking Mode Deep Thinking Streaming Messages Tool Streaming Output Function Calling Context Caching Structured Output Tools Web Search Stream Tool Call Agents GLM Slide/Poster Agent(beta) Translation Ag
GLM Coding Plan turns one year old today. To celebrate, we're giving every current subscriber a Reset Card. Use it to refill both your weekly and 5-hour quotas. Thanks for using GLM, helping shape it,
GLM Coding Plan turns one year old today. To celebrate, we're giving every current subscriber a Reset Card. Use it to refill both your weekly and 5-hour quotas. Thanks for using GLM, helping shape it, and pushing it to its limits. - Personal plan: https://t.co/3ut8Lnzhn4 -
ZCode now has 1 million users. As a thank-you to our community, we’ve reset usage limits for all GLM Coding Plan users. We’re also rolling out an update that helps turn long-horizon capabilities into
ZCode now has 1 million users. As a thank-you to our community, we’ve reset usage limits for all GLM Coding Plan users. We’re also rolling out an update that helps turn long-horizon capabilities into completed work: - More intelligence in real engineering workflows - 98% cache https://t.co/QT03v4lJVm
As models, contexts, and workloads grow, hidden assumptions in inference infrastructure can surface as output anomalies. Reliability requires more than throughput, latency, and availability. It also r
As models, contexts, and workloads grow, hidden assumptions in inference infrastructure can surface as output anomalies. Reliability requires more than throughput, latency, and availability. It also requires preserving the correctness of model state behind every generation.
After fixing correctness issues, we turned to the next bottleneck: Prefill throughput and GPU memory pressure in long-context Coding Agent serving. To address this, we introduced LayerSplit, a layer-w
After fixing correctness issues, we turned to the next bottleneck: Prefill throughput and GPU memory pressure in long-context Coding Agent serving. To address this, we introduced LayerSplit, a layer-wise KV Cache storage scheme. Instead of duplicating all layers on every GPU, https://t.co/OGptVovbtf
RT Cline: One week of @MistralAI's Devstral 2 in Cline. 6.52% diff-edit failure rate. On the diff-edit chart, that puts it ahead of GLM-4.6 (7.58%) an...
SkillX: Automatically Constructing Skill Knowledge Bases for Agents
Learning from experience is critical for building capable large language model (LLM) agents, yet prevailing self-evolving paradigms remain inefficient: agents learn in isolation, repeatedly rediscover similar behaviors from limited experience, resulting in redundant exploration and poor generalization. To address this problem, we propose SkillX, a fully automated framework for constructing a plug-and-play skill knowledge base that can be reused across agents and environments. SkillX operates through a fully automated pipeline built on three synergistic innovations: (i) Multi-Level Skills Design, which distills raw trajectories into three-tiered hierarchy of strategic plans, functional skills, and atomic skills; (ii) Iterative Skills Refinement, which automatically revises skills based on execution feedback to continuously improve library quality; and (iii) Exploratory Skills Expansion, which proactively generates and validates novel skills to expand coverage beyond seed training data. Using a strong backbone agent (GLM-4.6), we automatically build a reusable skill library and evaluate its transferability on challenging long-horizon, user-interactive benchmarks, including AppWorld, BFCL-v3, and τ^2-Bench. Experiments show that SkillKB consistently improves task success and execution efficiency when plugged into weaker base agents, highlighting the importance of structured, hierarchical experience representations for generalizable agent learning. Our code will be publicly available soon at https://github.com/zjunlp/SkillX.
Navigation GLM-4.6 Guides API Reference Coding Plan Released Notes Terms and Policy Help Center Get Started Quick Start Overview Pricing Core Parameters SDKs Guide Migrate to GLM-5.3 Models GLM-5.3-Flash/FlashX New GLM-5.3 Hot GLM-5.2 GLM-OCR GLM-ASR-2512 Capabilities Thinking Mode Deep Thinking Streaming Messages Tool Streaming Output Function Calling Context Caching Structured Output Tools Web Search Stream Tool Call Agents GLM Slide/Poster Agent(beta) Translation Agent Vid
Navigation Models GLM-5.3 Guides API Reference Coding Plan Released Notes Terms and Policy Help Center Get Started Quick Start Overview Pricing Core Parameters SDKs Guide Migrate to GLM-5.3 Models GLM-5.3-Flash/FlashX New GLM-5.3 Hot GLM-5.2 GLM-OCR GLM-ASR-2512 Capabilities Thinking Mode Deep Thinking Streaming Messages Tool Streaming Output Function Calling Context Caching Structured Output Tools Web Search Stream Tool Call Agents GLM Slide/Poster Agent(beta) Translation Ag