Incredibly cheap yet highly performant Chinese model, comparable to Deepseek R1 in performance on many metrics at 1/30th the cost.
Added Apr 15, 2025
Context Window
32.0K
Max Output
16.4K
Input Price (Auto)
$0.070/1M
Output Price (Auto)
$0.070/1M
Cache Read (Auto)
$0.035/1M
Capabilities
Benchmarks
Benchmarks
Performance metrics and benchmarks
No benchmark data is available yet for this model.
Providers
Auto routing is available for this model. Explicit provider selection is not available.
Loading provider options…
Related text models
Compare GLM Z1 Air with similar models from the same provider or model family.
GLM 4.5 Air
z-ai/GLM-4.5-AirGLM-4.5-Air is a 106B total / 12B active parameter model designed to unify frontier reasoning, coding, and agentic capabilities. On the SWE-bench Verified benchmark, it delivers the best performance at its scale with a competitive performance-to-cost ratio.
GLM 4 Air 0111
glm-4-air-0111MiniMax's flagship model with a 1M token context window
GLM 4.5 Air (Thinking)
z-ai/GLM-4.5-Air:thinkingGLM-4.5-Air with thinking mode enabled for enhanced reasoning capabilities. Shows step-by-step thought process.
GLM 5.3 TEE
TEE/glm-5.3GLM-5.3 is Z.AI's open-weight reasoning model for complex software engineering, autonomous agents, vulnerability research, and long-horizon tasks. This text-only deployment runs inside a Phala Trusted Execution Environment with Redpill attestation and signed completion receipts.
GLM 5.3 Flash TEE
TEE/glm-5.3-flashGLM-5.3 Flash is Z.AI's natively multimodal 320B MoE reasoning model with 18B active parameters, served by Phala inside a Trusted Execution Environment with Redpill attestation and signed completion receipts.
GLM 5.3 Flash
z-ai/glm-5.3-flashox-alpha out of stealth! GLM-5.3 Flash is Z.ai's first natively multimodal GLM-5 model, with 320B total parameters and just 18B active parameters for efficient coding, agentic work, and precise 1M-token context. Its hybrid sparse-and-linear attention architecture helps it outperform GLM-5.2 at one-tenth the price while approaching Claude Opus 4.8 on coding and agentic benchmarks.