GLM 5.3 Flashzhipu/glm-5-3-flash
GLM 5.3 Flash is a multimodal reasoning model for visual analysis, software development and agent tasks. The Flash architecture has 320B total parameters with 18B active per token; FlashX is the faster hosted serving option. Both accept visual inputs and generate text.
Availability and billing notes
Reviewed September 21, 2026 · check the linked provider documentation for later changes.
Reasoning is always enabled. The advertised serving speed is a vendor claim, not an independently measured latency result.
Prices below are USD rates from the international Z.ai API. Coding Plan allowances and mainland China pricing are separate.
Pricing across providers
| Provider | Input /1M | Output /1M | Blended /1M | Latency p50 | Format | Freshness | Action |
|---|---|---|---|---|---|---|---|
| Zhipu AI glm-5.3-flash International Z.ai API, USD; Coding Plan subscription usage is separate. | $0.15 Cached $0.03 | $0.5 | $0.237 | — | OpenAI-compatible | Needs 15d ago | Try → |
Affiliate disclosure: We may earn a commission from qualified signups. Pricing independence is enforced at the data layer — see our Editorial Independence Policy.
Works with
Point any of these clients at a hosting's base URL — they all speak at least one of this model's endpoint protocols (OPENAI_COMPATIBLE).
Capabilities
- reasoning
- code
- vision
- tool_calling
Code samples
Example using Zhipu AI — the cheapest hosting for this model as of last verification. Swap base_url and model to use a different provider from the matrix above.
from openai import OpenAI
client = OpenAI(
api_key="YOUR_API_KEY",
base_url="https://api.z.ai/api/paas/v4",
)
response = client.chat.completions.create(
model="glm-5.3-flash",
messages=[{"role": "user", "content": "Hello!"}],
)
print(response.choices[0].message.content)
Technical specs
- Context
- 1000K
- Max output
- —
- Parameters
- 320B
- Release
- —
- Training cutoff
- —
- License
- MIT
Similar models
Compare with
- GLM 5.3 Flash vs Qwen3.5 397B A17BComparison planned — not yet published
- GLM 5.3 Flash vs Step 3.7 FlashComparison planned — not yet published
- GLM 5.3 Flash vs GLM 5.3 FlashXComparison planned — not yet published
Frequently asked
How much does GLM 5.3 Flash cost?+−
How do I access GLM 5.3 Flash from outside China?+−
Is GLM 5.3 Flash open-source? Can I fine-tune it?+−
Is GLM 5.3 Flash OpenAI-compatible?+−
openai SDK client at the Provider's base_url and use the Provider's model name. See the Code Samples above for a copy-pasteable example.