Z.ai / Zhipu AIOpen-weight·1000K context·320B params·MIT

GLM 5.3 Flashzhipu/glm-5-3-flash

GLM 5.3 Flash is a multimodal reasoning model for visual analysis, software development and agent tasks. The Flash architecture has 320B total parameters with 18B active per token; FlashX is the faster hosted serving option. Both accept visual inputs and generate text.

Cheapest blended:$0.24 / 1M tokenson Zhipu AI · 1 provider listed

Availability and billing notes

Reviewed September 21, 2026 · check the linked provider documentation for later changes.

Reasoning is always enabled. The advertised serving speed is a vendor claim, not an independently measured latency result.

Prices below are USD rates from the international Z.ai API. Coding Plan allowances and mainland China pricing are separate.

Pricing across providers

Sort by:
ProviderInput /1MOutput /1MBlended /1MLatency p50FormatFreshnessAction
Zhipu AI
glm-5.3-flash

International Z.ai API, USD; Coding Plan subscription usage is separate.

$0.15
Cached $0.03
$0.5$0.237—OpenAI-compatibleNeeds 15d agoTry →

Affiliate disclosure: We may earn a commission from qualified signups. Pricing independence is enforced at the data layer — see our Editorial Independence Policy.

Works with

Point any of these clients at a hosting's base URL — they all speak at least one of this model's endpoint protocols (OPENAI_COMPATIBLE).

Capabilities

  • reasoning
  • code
  • vision
  • tool_calling

Code samples

Example using Zhipu AI — the cheapest hosting for this model as of last verification. Swap base_url and model to use a different provider from the matrix above.

from openai import OpenAI

client = OpenAI(
    api_key="YOUR_API_KEY",
    base_url="https://api.z.ai/api/paas/v4",
)

response = client.chat.completions.create(
    model="glm-5.3-flash",
    messages=[{"role": "user", "content": "Hello!"}],
)
print(response.choices[0].message.content)

Technical specs

Context
1000K
Max output
—
Parameters
320B
Release
—
Training cutoff
—
License
MIT

Similar models

Compare with

  • GLM 5.3 Flash vs Qwen3.5 397B A17B
    Comparison planned — not yet published
  • GLM 5.3 Flash vs Step 3.7 Flash
    Comparison planned — not yet published
  • GLM 5.3 Flash vs GLM 5.3 FlashX
    Comparison planned — not yet published

Frequently asked

How much does GLM 5.3 Flash cost?+
The cheapest public hosting is $0.24 per 1M blended tokens on Zhipu AI. 1 total providers are listed above with per-input / per-output / cached pricing.
How do I access GLM 5.3 Flash from outside China?+
Availability, account verification and payment methods vary by provider and region. Check the official documentation and the selected provider's access requirements before integrating. A directory listing does not guarantee global API access.
Is GLM 5.3 Flash open-source? Can I fine-tune it?+
GLM 5.3 Flash has published weights. License: MIT. Review the linked repository and license for deployment, fine-tuning and commercial-use conditions.
Is GLM 5.3 Flash OpenAI-compatible?+
At least one listed hosting exposes an OpenAI-compatible API, so you can point an existing openai SDK client at the Provider's base_url and use the Provider's model name. See the Code Samples above for a copy-pasteable example.
What's the maximum context window for GLM 5.3 Flash?+
The model supports up to 1,000,000 tokens of context (input + output). Some hosted versions may impose a smaller limit — check the "Context" column in the pricing matrix for each provider.