Fast GLM 5 Turbo variant from Z-AI for general chat, coding, and tool use.
Added Mar 15, 2026
Context Window
202.8K
Max Output
131.1K
Input Price (Auto)
$1.20/1M
Output Price (Auto)
$4.00/1M
Cache Read (Auto)
$0.24/1M
Capabilities
Benchmarks
Benchmarks
Performance metrics and benchmarks
Sourced from Artificial Analysis.
Intelligence Index
39.1
Reasoning
GPQA Diamond
Graduate-level scientific reasoning
84.7%
Better than 81% of models compared
HLE
Humanity's Last Exam
27.8%
Better than 82% of models compared
IFBench
Instruction-following benchmark
73.2%
Better than 90% of models compared
T²-Bench Telecom
Conversational AI agents in dual-control scenarios
98.5%
Better than 99% of models compared
AA-LCR
Long context reasoning evaluation
66.7%
Better than 72% of models compared
CritPt
Research-level physics reasoning
0.3%
Coding
SciCode
Python programming for scientific computing
43.6%
Better than 81% of models compared
Terminal-Bench Hard
Agentic coding and terminal use
33.3%
Better than 77% of models compared
Knowledge
AA-Omniscience Accuracy
Proportion of correctly answered questions
28.4%
AA-Omniscience Hallucination Rate
Rate of incorrect answers among non-correct responses
62.6%
Last updated Aug 16, 2026
Artificial AnalysisProviders
Auto routing is available for this model. Explicit provider selection is not available.
Loading provider options…
Related text models
Compare GLM 5 Turbo with similar models from the same provider or model family.
GLM 5V Turbo Thinking
z-ai/glm-5v-turbo:thinkingThinking-enabled GLM 5V Turbo for image, video, and text inputs. Uses the same multimodal foundation model with more deliberate vision-grounded analysis, planning, and tool use.
GLM 5V Turbo
z-ai/glm-5v-turboZ.ai's native multimodal agent model for vision-based coding and agent workflows. This is the standard non-thinking variant for image, video, and text inputs, tuned for perceive-plan-execute loops, complex coding, and tool-driven task execution.
GLM 4.6 Turbo
z-ai/GLM-4.6-turboFast variant of GLM 4.6 for general chat, coding, and analysis with improved latency and strong reasoning.
GLM 4.6 Turbo (Thinking)
z-ai/GLM-4.6-turbo:thinkingGLM 4.6 Turbo with thinking mode enabled for enhanced reasoning; shows internal reasoning and supports long context.
GLM 5.3 Flash
z-ai/glm-5.3-flashox-alpha out of stealth! GLM-5.3 Flash is Z.ai's first natively multimodal GLM-5 model, with 320B total parameters and just 18B active parameters for efficient coding, agentic work, and precise 1M-token context. Its hybrid sparse-and-linear attention architecture helps it outperform GLM-5.2 at one-tenth the price while approaching Claude Opus 4.8 on coding and agentic benchmarks.
GLM 5.3
z-ai/glm-5.3GLM-5.3 for long-horizon autonomous coding and engineering workflows. This variant defaults to the model's low reasoning tier for faster responses.