ox-alpha out of stealth! GLM-5.3 Flash is Z.ai's first natively multimodal GLM-5 model, with 320B total parameters and just 18B active parameters for efficient coding, agentic work, and precise 1M-token context. Its hybrid sparse-and-linear attention architecture helps it outperform GLM-5.2 at one-tenth the price while approaching Claude Opus 4.8 on coding and agentic benchmarks.
Added Aug 26, 2026
Model weightsContext Window
1.0M
Max Output
131.1K
Input Price (Auto)
$0.075/1M
Output Price (Auto)
$0.25/1M
Cache Read (Auto)
$0.015/1M
Capabilities
Benchmarks
Benchmarks
Performance metrics and benchmarks
Sourced from LMArena.
Arena Score
1469.4
Overall Rank
#38 / 395
Votes
2,424
Confidence Interval
1457.4 - 1481.4
Category Scores
Coding
#11 / 390
663 votes
1530.6
Longer Query
#45 / 373
1,135 votes
1474.2
Creative Writing
#58 / 393
541 votes
1433.3
Instruction Following
#31 / 395
860 votes
1465.1
Hard Prompts
#30 / 395
1,607 votes
1495.0
Additional Categories13
Expert
#15 / 345
238 votes
1514.2
Hard Prompts English
#30 / 393
577 votes
1497.7
Non English
#34 / 395
1,521 votes
1459.4
Industry Software And It Services
#35 / 395
935 votes
1503.4
Exclude Ties
#38 / 395
1,771 votes
1474.1
Russian
#43 / 359
282 votes
1466.4
English
#45 / 395
903 votes
1471.8
Industry Entertainment And Sports And Media
#46 / 393
614 votes
1436.8
Industry Writing And Literature And Language
#57 / 394
660 votes
1443.4
Industry Business And Management And Financial Operations
#69 / 388
436 votes
1446.7
Multi Turn
#71 / 393
356 votes
1452.8
Industry Life And Physical And Social Science
#74 / 393
408 votes
1461.5
Industry Legal And Government
#76 / 368
229 votes
1455.6
Published 2026-08-27 · Matched as glm-5.3-flash
LMArena DatasetProviders
Choose explicit providers for this model. Auto routing remains available as the default option.
Loading provider options…
Related text models
Compare GLM 5.3 Flash with similar models from the same provider or model family.
GLM 5.3 Flash Uncensored
z-ai/glm-5.3-flash-uncensoredGLM 5.3 Flash Uncensored is an FP8 uncensored fine-tune of the efficient 320B mixture-of-experts reasoning model with restored vision support, built for unrestricted chat, creative writing, coding, agentic work, tool use, and long-context tasks.
GLM 4.7 Flash
z-ai/glm-4.7-flashGLM-4.7-Flash is a lightweight 30B model optimized for coding and agentic tasks. Balances high performance with efficiency.
GLM 4.7 Flash Original
z-ai/glm-4.7-flash-originalGLM-4.7-Flash is a lightweight 30B model optimized for coding and agentic tasks. Balances high performance with efficiency, perfect for local deployment.
GLM 4.7 Flash Original Thinking
z-ai/glm-4.7-flash-original:thinkingGLM-4.7-Flash with extended thinking capabilities for complex reasoning. Lightweight 30B model optimized for coding and agentic tasks.
GLM 4.7 Flash Thinking
z-ai/glm-4.7-flash:thinkingGLM-4.7-Flash with extended thinking capabilities for complex reasoning. Lightweight 30B model optimized for coding and agentic tasks.
GLM 4.6V Flash
z-ai/glm-4.6v-flash-originalGLM-4.6V-Flash (9B), a lightweight model optimized for local deployment and low-latency applications. Scales context window to 128k tokens and achieves SoTA performance in visual understanding among similar-scale models.