Z.ai / Zhipu AIWeights not listed·1000K context·320B params

GLM 5.3 FlashXzhipu/glm-5-3-flashx

GLM 5.3 FlashX is a multimodal reasoning model for visual analysis, software development and agent tasks. The Flash architecture has 320B total parameters with 18B active per token; FlashX is the faster hosted serving option. Both accept visual inputs and generate text.

Cheapest blended:$0.59 / 1M tokenson Zhipu AI · 1 provider listed

Availability and billing notes

Reviewed September 21, 2026 · check the linked provider documentation for later changes.

Reasoning is always enabled. The advertised serving speed is a vendor claim, not an independently measured latency result.

Prices below are USD rates from the international Z.ai API. Coding Plan allowances and mainland China pricing are separate.

Pricing across providers

Sort by:
ProviderInput /1MOutput /1MBlended /1MLatency p50FormatFreshnessAction
Zhipu AI
glm-5.3-flashx

International Z.ai API, USD; Coding Plan subscription usage is separate.

$0.37
Cached $0.075
$1.25$0.59—OpenAI-compatibleNeeds 15d agoTry →

Affiliate disclosure: We may earn a commission from qualified signups. Pricing independence is enforced at the data layer — see our Editorial Independence Policy.

Works with

Point any of these clients at a hosting's base URL — they all speak at least one of this model's endpoint protocols (OPENAI_COMPATIBLE).

Capabilities

  • reasoning
  • code
  • vision
  • tool_calling

Code samples

Example using Zhipu AI — the cheapest hosting for this model as of last verification. Swap base_url and model to use a different provider from the matrix above.

from openai import OpenAI

client = OpenAI(
    api_key="YOUR_API_KEY",
    base_url="https://api.z.ai/api/paas/v4",
)

response = client.chat.completions.create(
    model="glm-5.3-flashx",
    messages=[{"role": "user", "content": "Hello!"}],
)
print(response.choices[0].message.content)

Technical specs

Context
1000K
Max output
—
Parameters
320B
Release
—
Training cutoff
—
License
—

Similar models

Compare with

  • GLM 5.3 FlashX vs GLM 5.3 Flash
    Comparison planned — not yet published
  • GLM 5.3 FlashX vs Qwen3.5 397B A17B
    Comparison planned — not yet published
  • GLM 5.3 FlashX vs Step 3.7 Flash
    Comparison planned — not yet published

Frequently asked

How much does GLM 5.3 FlashX cost?+
The cheapest public hosting is $0.59 per 1M blended tokens on Zhipu AI. 1 total providers are listed above with per-input / per-output / cached pricing.
How do I access GLM 5.3 FlashX from outside China?+
Availability, account verification and payment methods vary by provider and region. Check the official documentation and the selected provider's access requirements before integrating. A directory listing does not guarantee global API access.
Is GLM 5.3 FlashX open-source?+
Downloadable weights are not listed in this directory for GLM 5.3 FlashX. See the official documentation and availability notes for API, subscription or planned weight-release details.
Is GLM 5.3 FlashX OpenAI-compatible?+
At least one listed hosting exposes an OpenAI-compatible API, so you can point an existing openai SDK client at the Provider's base_url and use the Provider's model name. See the Code Samples above for a copy-pasteable example.
What's the maximum context window for GLM 5.3 FlashX?+
The model supports up to 1,000,000 tokens of context (input + output). Some hosted versions may impose a smaller limit — check the "Context" column in the pricing matrix for each provider.