We're introducing GLM-5. 2, our latest flagship model for long-horizon tasks. It marks a substantial leap in long-horizon task capability over its predecessor GLM-5. 1 and, for the first time, delivers that capability on a solid 1M-token context.
Running this yourself: likely needs a high-memory cloud gpu.
Model updates refreshed3h agoSep 28, 2026news + changelog
Live access is not confirmed for this model. No current purchase price is advertised.
Related provider subscriptions
No current subscription pricing is tracked for this model.
Confirm this specific model, usage limits, and billing terms with the provider. A subscription does not automatically include API credits.
No login needed to compare. Prices are in USD; provider charges are separate from AI Market Cap plans. Context length, caching, tools, taxes, and regional terms can change the final cost. Open weights do not mean free hosting.
60.8
Quality Score
---
Arena ELO
753B
Parameters
1M
Context
Evidence profile
How complete is this record?
This measures the amount of verifiable public evidence we have, not how capable the model is. A missing field means it has not been verified yet, not that its value is zero.
Epoch AI, Data on AI Models. Used under CC BY with attribution.
Launches
1
high
Benchmarks
14
high
API
2
medium
Research
4
low
General
5
low
What Changed Recently
Recent launch, pricing, benchmark, and API signals linked to this model or its provider.
LaunchesZ.ai1mo ago
ZCode now has 1 million users. As a thank-you to our community, we’ve reset usage limits for all GLM Coding Plan users. We’re also rolling out an update that helps turn long-horizon capabilities into
ZCode now has 1 million users. As a thank-you to our community, we’ve reset usage limits for all GLM Coding Plan users. We’re also rolling out an update that helps turn long-horizon capabilities into completed work: - More intelligence in real engineering workflows - 98% cache https://t.co/QT03v4lJVm
Navigation Models GLM-5.3-Flash/FlashX Guides API Reference Coding Plan Released Notes Terms and Policy Help Center Get Started Quick Start Overview Pricing Core Parameters SDKs Guide Migrate to GLM-5.3 Models GLM-5.3-Flash/FlashX New GLM-5.3 Hot GLM-5.2 GLM-OCR GLM-ASR-2512 Capabilities Thinking Mode Deep Thinking Streaming Messages Tool Streaming Output Function Calling Context Caching Structured Output Tools Web Search Stream Tool Call Agents GLM Slide/Poster Agent(beta) T
Navigation Models GLM-ASR-2512 Guides API Reference Coding Plan Released Notes Terms and Policy Help Center Get Started Quick Start Overview Pricing Core Parameters SDKs Guide Migrate to GLM-5.3 Models GLM-5.3-Flash/FlashX New GLM-5.3 Hot GLM-5.2 GLM-OCR GLM-ASR-2512 Capabilities Thinking Mode Deep Thinking Streaming Messages Tool Streaming Output Function Calling Context Caching Structured Output Tools Web Search Stream Tool Call Agents GLM Slide/Poster Agent(beta) Translati
Navigation Models GLM-5.3-Flash/FlashX Guides API Reference Coding Plan Released Notes Terms and Policy Help Center Get Started Quick Start Overview Pricing Core Parameters SDKs Guide Migrate to GLM-5.3 Models GLM-5.3-Flash/FlashX New GLM-5.3 Hot GLM-5.2 GLM-OCR GLM-ASR-2512 Capabilities Thinking Mode Deep Thinking Streaming Messages Tool Streaming Output Function Calling Context Caching Structured Output Tools Web Search Stream Tool Call Agents GLM Slide/Poster Agent(beta) T
Navigation Models GLM-5.3-Flash/FlashX Guides API Reference Coding Plan Released Notes Terms and Policy Help Center Get Started Quick Start Overview Pricing Core Parameters SDKs Guide Migrate to GLM-5.3 Models GLM-5.3-Flash/FlashX New GLM-5.3 Hot GLM-5.2 GLM-OCR GLM-ASR-2512 Capabilities Thinking Mode Deep Thinking Streaming Messages Tool Streaming Output Function Calling Context Caching Structured Output Tools Web Search Stream Tool Call Agents GLM Slide/Poster Agent(beta) T
GLM Coding Plan turns one year old today. To celebrate, we're giving every current subscriber a Reset Card. Use it to refill both your weekly and 5-hour quotas. Thanks for using GLM, helping shape it,
GLM Coding Plan turns one year old today. To celebrate, we're giving every current subscriber a Reset Card. Use it to refill both your weekly and 5-hour quotas. Thanks for using GLM, helping shape it, and pushing it to its limits. - Personal plan: https://t.co/3ut8Lnzhn4 -
ZCode now has 1 million users. As a thank-you to our community, we’ve reset usage limits for all GLM Coding Plan users. We’re also rolling out an update that helps turn long-horizon capabilities into
ZCode now has 1 million users. As a thank-you to our community, we’ve reset usage limits for all GLM Coding Plan users. We’re also rolling out an update that helps turn long-horizon capabilities into completed work: - More intelligence in real engineering workflows - 98% cache https://t.co/QT03v4lJVm
As models, contexts, and workloads grow, hidden assumptions in inference infrastructure can surface as output anomalies. Reliability requires more than throughput, latency, and availability. It also r
As models, contexts, and workloads grow, hidden assumptions in inference infrastructure can surface as output anomalies. Reliability requires more than throughput, latency, and availability. It also requires preserving the correctness of model state behind every generation.
After fixing correctness issues, we turned to the next bottleneck: Prefill throughput and GPU memory pressure in long-context Coding Agent serving. To address this, we introduced LayerSplit, a layer-w
After fixing correctness issues, we turned to the next bottleneck: Prefill throughput and GPU memory pressure in long-context Coding Agent serving. To address this, we introduced LayerSplit, a layer-wise KV Cache storage scheme. Instead of duplicating all layers on every GPU, https://t.co/OGptVovbtf
We present Miles v0.1, a full-stack, production-ready system for frontier post-training. Building upon the clean design of slime, Miles designs each stage of the reinforcement-learning (RL) training loop around a single principle: components should be verified, clean, and customizable. With accuracy, efficiency, reliability, and scalability as first-class goals, Miles aims to make frontier-scale RL accessible to researchers and enterprises alike. This report walks through the system end to end: rollout engines built on SGLang, a trainer with a choice of two backends (NVIDIA Megatron-LM and PyTorch FSDP), and three weight-synchronization transports for different deployment topologies. Beyond full-parameter RL, Miles also supports LoRA RL, on-policy distillation, supervised fine-tuning, and true-on-policy rollout-training alignment, and extends the same architecture to diffusion models. We close with an end-to-end case study: fully asynchronous agentic RL on a GLM-5.2 744B-A40B model over terminal-use coding tasks, running on 64 NVIDIA GB300 GPUs with a median step time of 263 seconds over the first 30 measured steps. Miles is open-sourced at https://github.com/radixark/miles, with the project website at https://miles.radixark.com.
MindForge: Teaching Small Language Models Whole-Life-Cycle Software Engineering via Source-Free Program Synthesis
Coding agents have made substantial progress on software engineering tasks that modify existing codebases, including bug fixing and feature implementation. However, constructing a complete program from scratch remains a major challenge: even the frontier models evaluated on ProgramBench fully resolve fewer than 1% of tasks. One obstacle is the lack of scalable training environments for this from-scratch setting, spanning the whole software engineering life cycle, as existing environment-construction frameworks focus only on a single phase in software development. To address this gap, we introduce MindForge, an automated pipeline that converts open-source command-line programs into source-free environments that expose only a compiled reference executable and its documentation. Using MindForge, we construct training environments from repositories disjoint from those in ProgramBench, and curate a high-quality data recipe consisting of program synthesis trajectories using GLM-5.2 as the teacher agent. Fine-tuning Qwen3.6-27B on these trajectories increases its ProgramBench average test pass rate from 37.98% to 49.51%, achieving performance comparable to substantially larger frontier models. Moreover, the fine-tuned model consistently improves over the base model across all seven unseen software engineering benchmarks, spanning long-horizon repository generation and translation, bug fixing, feature implementation, and cross-language issue resolution, with absolute gains of 31.00 points on RepoZero-C2Rust, 14.16 on DeepSWE, 10.70/4.56 on NL2Repo-Bench (with/without tests), 5.04 on SWE-bench Verified, 5.93 on SWE-bench Pro, 5.22 on SWE-bench Multilingual, and 4.94 on FeatBench.
LongStraw: Long-Context RL Beyond 2M Tokens under a Fixed GPU Budget
A growing gap separates inference context lengths from RL post-training: inference systems are approaching million-token contexts, while post-training workloads often remain at 256K tokens or below and rely on length generalization at deployment. The gap is especially important for AI agents, whose observations, tool outputs, documents, and prior decisions accumulate over long trajectories. LongStraw is an architecture-aware execution stack for million-token RL post-training under a fixed GPU budget, instantiated with Group Relative Policy Optimization (GRPO). It evaluates the shared prompt without autograd, retains only model-specific state needed by later tokens, and replays short response branches one at a time, reducing the live training graph at the cost of additional replay time. We implement it for the hybrid recurrent and full-attention Qwen3.6-27B and the compressed-attention mixture-of-experts GLM-5.2. On eight H20 GPUs, LongStraw completes grouped Qwen scoring and response backward at 2.1M positions for groups of 2 and 8; increasing the group size adds only 0.21 GB of peak allocated memory, while a separate stress test reaches 4.46M positions. On 32 H20 GPUs, we validate the end-to-end LongStraw execution path for a 2.1M-token prompt across all 78 layers of GLM-5.2. These experiments establish execution capacity rather than complete training correctness because the captured prompt state is detached and some distributed forward and gradient composition paths remain incomplete.
Single-Rollout Asynchronous Optimization for Agentic Reinforcement Learning
Reinforcement learning (RL) is becoming increasingly important for post-training large language models (LLMs). Previous RL pipelines for LLMs were mostly synchronous and batch-interleaved, which is inefficient for long-horizon agentic tasks. Recently, asynchronous RL has emerged as a more efficient alternative by updating the model as rollouts arrive. However, existing asynchronous RL systems often emphasize throughput, while leaving training stability and task effectiveness largely underexplored. For example, a key challenge is that group-wise sampling in the widely-used GRPO framework does not naturally fit asynchronous agentic training. In this paper, we present Single-rollout Asynchronous Optimization (SAO) to address the stability and off-policy challenges in asynchronous RL. To reduce off-policy effects and improve generalization, we replace group-wise sampling with single-rollout sampling, that is, using one rollout per prompt. We further improve this single-rollout strategy with practical value-model training designs. To improve optimization stability, we introduce a strict double-side token-level clipping strategy. SAO is able to train stably for one thousand steps and consistently outperform GRPO and its variants on agentic coding and reasoning benchmarks, such as SWE-Bench Verified, BeyondAIME, and IMOAnswerBench. We also demonstrate that single-rollout RL is particularly effective in a simulated online learning setting, where the model must adapt to changing evolving environments. To this end, SAO is successfully deployed in the agentic RL pipeline for training the open GLM-5.2 model (750B-A40B).
Navigation Models GLM-5.3-Flash/FlashX Guides API Reference Coding Plan Released Notes Terms and Policy Help Center Get Started Quick Start Overview Pricing Core Parameters SDKs Guide Migrate to GLM-5.3 Models GLM-5.3-Flash/FlashX New GLM-5.3 Hot GLM-5.2 GLM-OCR GLM-ASR-2512 Capabilities Thinking Mode Deep Thinking Streaming Messages Tool Streaming Output Function Calling Context Caching Structured Output Tools Web Search Stream Tool Call Agents GLM Slide/Poster Agent(beta) T
Navigation Models GLM-ASR-2512 Guides API Reference Coding Plan Released Notes Terms and Policy Help Center Get Started Quick Start Overview Pricing Core Parameters SDKs Guide Migrate to GLM-5.3 Models GLM-5.3-Flash/FlashX New GLM-5.3 Hot GLM-5.2 GLM-OCR GLM-ASR-2512 Capabilities Thinking Mode Deep Thinking Streaming Messages Tool Streaming Output Function Calling Context Caching Structured Output Tools Web Search Stream Tool Call Agents GLM Slide/Poster Agent(beta) Translati
Navigation Models GLM-5.3-Flash/FlashX Guides API Reference Coding Plan Released Notes Terms and Policy Help Center Get Started Quick Start Overview Pricing Core Parameters SDKs Guide Migrate to GLM-5.3 Models GLM-5.3-Flash/FlashX New GLM-5.3 Hot GLM-5.2 GLM-OCR GLM-ASR-2512 Capabilities Thinking Mode Deep Thinking Streaming Messages Tool Streaming Output Function Calling Context Caching Structured Output Tools Web Search Stream Tool Call Agents GLM Slide/Poster Agent(beta) T
Navigation Models GLM-5.3-Flash/FlashX Guides API Reference Coding Plan Released Notes Terms and Policy Help Center Get Started Quick Start Overview Pricing Core Parameters SDKs Guide Migrate to GLM-5.3 Models GLM-5.3-Flash/FlashX New GLM-5.3 Hot GLM-5.2 GLM-OCR GLM-ASR-2512 Capabilities Thinking Mode Deep Thinking Streaming Messages Tool Streaming Output Function Calling Context Caching Structured Output Tools Web Search Stream Tool Call Agents GLM Slide/Poster Agent(beta) T
Navigation Models GLM-5.3-Flash/FlashX Guides API Reference Coding Plan Released Notes Terms and Policy Help Center Get Started Quick Start Overview Pricing Core Parameters SDKs Guide Migrate to GLM-5.3 Models GLM-5.3-Flash/FlashX New GLM-5.3 Hot GLM-5.2 GLM-OCR GLM-ASR-2512 Capabilities Thinking Mode Deep Thinking Streaming Messages Tool Streaming Output Function Calling Context Caching Structured Output Tools Web Search Stream Tool Call Agents GLM Slide/Poster Agent(beta) T
Navigation Models GLM-5.3-Flash/FlashX Guides API Reference Coding Plan Released Notes Terms and Policy Help Center Get Started Quick Start Overview Pricing Core Parameters SDKs Guide Migrate to GLM-5.3 Models GLM-5.3-Flash/FlashX New GLM-5.3 Hot GLM-5.2 GLM-OCR GLM-ASR-2512 Capabilities Thinking Mode Deep Thinking Streaming Messages Tool Streaming Output Function Calling Context Caching Structured Output Tools Web Search Stream Tool Call Agents GLM Slide/Poster Agent(beta) T
Navigation Models GLM-OCR Guides API Reference Coding Plan Released Notes Terms and Policy Help Center Get Started Quick Start Overview Pricing Core Parameters SDKs Guide Migrate to GLM-5.3 Models GLM-5.3-Flash/FlashX New GLM-5.3 Hot GLM-5.2 GLM-OCR GLM-ASR-2512 Capabilities Thinking Mode Deep Thinking Streaming Messages Tool Streaming Output Function Calling Context Caching Structured Output Tools Web Search Stream Tool Call Agents GLM Slide/Poster Agent(beta) Translation Ag
Navigation Models GLM-5.3-Flash/FlashX Guides API Reference Coding Plan Released Notes Terms and Policy Help Center Get Started Quick Start Overview Pricing Core Parameters SDKs Guide Migrate to GLM-5.3 Models GLM-5.3-Flash/FlashX New GLM-5.3 Hot GLM-5.2 GLM-OCR GLM-ASR-2512 Capabilities Thinking Mode Deep Thinking Streaming Messages Tool Streaming Output Function Calling Context Caching Structured Output Tools Web Search Stream Tool Call Agents GLM Slide/Poster Agent(beta) T
Navigation GLM-5V-Turbo Guides API Reference Coding Plan Released Notes Terms and Policy Help Center Get Started Quick Start Overview Pricing Core Parameters SDKs Guide Migrate to GLM-5.3 Models GLM-5.3-Flash/FlashX New GLM-5.3 Hot GLM-5.2 GLM-OCR GLM-ASR-2512 Capabilities Thinking Mode Deep Thinking Streaming Messages Tool Streaming Output Function Calling Context Caching Structured Output Tools Web Search Stream Tool Call Agents GLM Slide/Poster Agent(beta) Translation Agen
Navigation Models GLM-5.3 Guides API Reference Coding Plan Released Notes Terms and Policy Help Center Get Started Quick Start Overview Pricing Core Parameters SDKs Guide Migrate to GLM-5.3 Models GLM-5.3-Flash/FlashX New GLM-5.3 Hot GLM-5.2 GLM-OCR GLM-ASR-2512 Capabilities Thinking Mode Deep Thinking Streaming Messages Tool Streaming Output Function Calling Context Caching Structured Output Tools Web Search Stream Tool Call Agents GLM Slide/Poster Agent(beta) Translation Ag
voxtral-mini-transcribe-realtime-2602 +1 Speed Performance Modalities Context i -- Price i $ 0.006 /Min Speed Performance Modalities Context -- Price $ 0.006 /Min FEATURES WEIGHTS Features Transcriptions /v1/audio/transcriptions Other Models Other Models Z.ai GLM 5.3 v 5.3 Z.ai GLM 5.2 v 5.2 Shieldstral 1.0 v 1.0 WHY MISTRAL About us Our customers Careers Contact us EXPLORE AI Solutions Partners Research DOCUMENTATION Documentation Ambassadors Cookbooks BUILD Studio Vibe Mist