GLM 4.5 Air Pricing (2026)

GLM-4.5-Air is Zhipu's budget tier, and the cheapest paid way to reach it is OpenRouter at $0.13 input / $0.85 output per million tokens, undercutting…

Fan Chuanyu's profile

Written by Fan Chuanyu

7 min read

GLM-4.5-Air is Zhipu's budget tier, and the cheapest paid way to reach it is OpenRouter at $0.13 input / $0.85 output per million tokens, undercutting Zhipu's own official $0.20 / $1.10 list price by roughly a third. GLM-4.5-Air is the small-model rung of the GLM family: it trades a little reasoning depth for a rate card that is about 40% cheaper than GLM-4.6 on output, which is where most chat and agent bills concentrate. If you are routing through OpenRouter rather than a native Zhipu key, that is the number you will actually pay, and we measured a real call to prove it fires.

GLM-4.5-Air API pricing (verified 2026-06)

GLM-4.5-Air is a budget-tier large language model that delivers near-mid-tier quality at small-model token rates. The price you see depends entirely on how you reach it, and the two paths diverge by about a third on input.

According to Z.ai pricing, GLM-4.5-Air on the official Zhipu international platform lists at $0.20 input / $1.10 output per million tokens. Routed through OpenRouter, the same model is cataloged at $0.13 / $0.85. That is the same OpenRouter-undercuts-official pattern we have seen across the Chinese-model field, and it is the reason this article leads with the routed number: it is what your invoice reflects when the request leaves through a third party. We have not yet re-verified the official figure against a native Zhipu key, so treat the $0.20 / $1.10 row as sourced research pending that check.

Path to GLM-4.5-AirInput ($/1M)Output ($/1M)Notes
Measured via OpenRouter$0.13$0.85Live catalog, billed on real calls
Official Z.ai list$0.20$1.10needs native re-verify
GLM-4.6 via OpenRouter$0.43$1.74Step up in tier and price

The practical upshot: on output, the lever that dominates most production bills, Air via OpenRouter is roughly half the cost of GLM-4.6. That gap is the whole reason the budget tier exists.

First-hand: what one Air call actually billed

Documentation tells you the rate. We wanted the receipt. When we called GLM-4.5-Air through OpenRouter on 2026-06-26, a single general-purpose request of 32 input tokens and 200 output tokens billed $0.00017374 and returned in 8.07 seconds. The output hit our 200-token cap, so treat that completion length as capped rather than the model's natural stopping point.

That number matters less as an absolute and more as a ratio. We ran the identical prompt against GLM-4.6 through the same route in the same session: 32 in, 200 out, $0.00036376, 8.58 seconds. So on a matched call, Air cost 52% less and finished about half a second faster. The latency edge is small and within noise, but the cost edge is structural and repeats on every call. For a high-volume chat or classification workload, that 52% compounds into the difference between a comfortable budget and an alarming one.

Both numbers are OpenRouter-routed and live-billed; they are not list-price arithmetic. The full run set lives in our evidence pack.

When Air beats GLM-4.6 on cost-per-quality

The honest framing is cost-per-acceptable-answer, not cost-per-token. Air wins when the task tolerates a slightly shallower model: routing, summarization, classification, extraction, first-draft generation, and most high-turn chat. On those, paying GLM-4.6 rates buys reasoning headroom you never use, and the 52% we measured per matched call is pure waste.

GLM-4.6 earns its premium on the harder end: multi-step reasoning, longer-context synthesis, and the coding tasks where one extra correct token saves a retry. The decision rule is blunt. Choose GLM-4.5-Air when your eval shows Air clears your quality bar, because then you are simply overpaying with anything heavier. Move up to GLM-4.6 when Air's miss rate forces enough retries that the cheaper per-call rate stops being cheaper per-success.

One tier sits below Air on price but carries a sharp caveat. According to Zhipu open platform docs, the GLM-4.7-Flash and GLM-4.5-Flash models are offered free on the Zhipu international platform, a genuine zero-cost tier that OpenRouter does not match. That free tier is an official-Zhipu fact, not an OpenRouter price. But cheaper is not always faster: our Flash call through OpenRouter was the slowest in the whole set at 26.24 seconds, a reminder that "flash" naming does not guarantee flash latency. If your workload is latency-sensitive, Air's 8-second response may be worth its modest cost over a free tier that stalls.

How GLM-4.5-Air compares to DeepSeek

For non-Zhipu buyers weighing budget options, DeepSeek is the obvious cross-check, and here the comparison flips. According to DeepSeek API Docs, DeepSeek V4-Flash bills $0.14 input / $0.28 output per million tokens on its official endpoint. On input, DeepSeek-official and Air-via-OpenRouter are nearly level at $0.14 versus $0.13. On output, DeepSeek's $0.28 is a third of Air's $0.85, a decisive gap.

Model (cheapest path)Input ($/1M)Output ($/1M)
GLM-4.5-Air via OpenRouter$0.13$0.85
DeepSeek V4-Flash official$0.14$0.28

Note the asymmetry in the comparison itself: the GLM number is OpenRouter-routed while the DeepSeek number is the official first-party rate, so this is not a same-platform contest. The takeaway is still clear. If raw output cost is your only axis, DeepSeek V4-Flash is the cheaper budget model. You would pick GLM-4.5-Air instead when you specifically want Zhipu's model behavior, its Chinese-language strengths, or when an eval shows Air's outputs clear your bar at a token count low enough to close the per-token gap. For the deeper DeepSeek rate breakdown, see the DeepSeek API pricing hub.

FAQ

How much does GLM-4.5-Air cost per million tokens? Measured via OpenRouter, GLM-4.5-Air is $0.13 input / $0.85 output per million tokens. Zhipu's official international list price is higher at $0.20 / $1.10, pending native-key re-verification. The OpenRouter route is the cheaper paid path.

Is GLM-4.5-Air cheaper than GLM-4.6? Yes, substantially. On a matched live call we ran 2026-06-26, Air billed $0.00017374 versus GLM-4.6's $0.00036376 for the same 32-in / 200-out prompt through OpenRouter, making Air 52% cheaper on that call and roughly half the price on output rates.

Is there a free way to use a GLM model? Yes, but not Air. According to Zhipu, the GLM-4.7-Flash and GLM-4.5-Flash tiers are free on the Zhipu international platform. Our Flash call was slow, though, at 26.24 seconds, so the free tier trades latency for cost. GLM-4.5-Air is the cheapest paid tier with faster response.

Should I pick GLM-4.5-Air or DeepSeek V4-Flash for budget work? On output cost, DeepSeek V4-Flash ($0.28/1M) clearly beats Air ($0.85/1M via OpenRouter), with input near level. Choose Air when you specifically want Zhipu's model behavior; choose DeepSeek when output cost is your only axis.


This is part of the GLM API pricing hub, which covers every GLM tier side by side.

Author: Kevin Fan, Customer Success Manager at China LLM Directory, specializing in Chinese LLM ecosystem pricing and API cost benchmarking. Last verified: 2026-06-26.

<!-- METADATA { "title": "GLM-4.5-Air API Pricing: Budget Tier (2026)", "slug": "glm-4.5-air-pricing", "meta_description": "GLM-4.5-Air costs $0.13/$0.85 per 1M via OpenRouter vs Zhipu's official $0.20/$1.10. Live test: 32/200 tokens billed $0.00017374 in 8.07s. When Air beats 4.6.", "focus_keyword": "glm-4.5-air pricing", "secondary_keywords": ["glm-4.5-air api cost", "glm air token price", "glm budget tier pricing", "glm-4.5-air vs glm-4.6 cost"], "tags": ["GLM", "Zhipu", "API Pricing", "Budget Models"], "category": "Pricing", "cluster_id": "glm-api-pricing", "cluster_role": "micro", "hub_slug": "glm-api-pricing", "evidence_file": "clients/china-llm-aggregator/articles/glm-api-pricing-evidence.json", "needs_native_reverify": true, "verified_until": "2026-09-24", "author_name": "Kevin Fan", "author_title": "Customer Success Manager", "author_linkedin": "", "author_expertise": ["Chinese LLM ecosystem", "AI infrastructure pricing", "model benchmarking", "cross-border AI compliance"], "faq_pairs": [ {"q": "How much does GLM-4.5-Air cost per million tokens?", "a": "Measured via OpenRouter, GLM-4.5-Air is $0.13 input / $0.85 output per million tokens. Zhipu's official international list price is higher at $0.20 / $1.10, pending native-key re-verification. The OpenRouter route is the cheaper paid path."}, {"q": "Is GLM-4.5-Air cheaper than GLM-4.6?", "a": "Yes, substantially. On a matched live call we ran 2026-06-26, Air billed $0.00017374 versus GLM-4.6's $0.00036376 for the same 32-in / 200-out prompt through OpenRouter, making Air 52% cheaper on that call and roughly half the price on output rates."}, {"q": "Is there a free way to use a GLM model?", "a": "Yes, but not Air. According to Zhipu, the GLM-4.7-Flash and GLM-4.5-Flash tiers are free on the Zhipu international platform. Our Flash call was slow, though, at 26.24 seconds, so the free tier trades latency for cost. GLM-4.5-Air is the cheapest paid tier with faster response."}, {"q": "Should I pick GLM-4.5-Air or DeepSeek V4-Flash for budget work?", "a": "On output cost, DeepSeek V4-Flash ($0.28/1M) clearly beats Air ($0.85/1M via OpenRouter), with input near level. Choose Air when you specifically want Zhipu's model behavior; choose DeepSeek when output cost is your only axis."} ], "external_links_used": [ {"url": "https://z.ai/", "source_name": "Z.ai", "claim": "GLM-4.5-Air official list price $0.20 in / $1.10 out per 1M (needs native re-verify)"}, {"url": "https://bigmodel.cn/", "source_name": "Zhipu open platform docs", "claim": "GLM-4.7-Flash and GLM-4.5-Flash free tier on Zhipu international platform"}, {"url": "https://api-docs.deepseek.com/quick_start/pricing/", "source_name": "DeepSeek API Docs", "claim": "DeepSeek V4-Flash $0.14 in / $0.28 out per 1M official endpoint"} ], "internal_links_used": [ {"url": "/blog/glm-api-pricing/", "anchor_text": "GLM API pricing hub", "type": "hub"}, {"url": "/blog/deepseek-api-pricing/", "anchor_text": "DeepSeek API pricing hub", "type": "cross-cluster"} ], "first_hand_evidence": { "source": "glm-api-pricing-evidence.json run label glm-4.5-air_general (+ glm-4.6_general for matched comparison)", "measured": "GLM-4.5-Air via OpenRouter: 32 in / 200 out, $0.00017374, 8.07s. Matched GLM-4.6: 32 in / 200 out, $0.00036376, 8.58s. Flash call slowest at 26.24s.", "captured": "2026-06-26", "routing": "GLM tiers OpenRouter-routed and live-billed; DeepSeek anchor is official api.deepseek.com" }, "images_status": "spec-only (not generated; FAL_API_KEY unset)", "images": [ {"position": "featured", "type": "generated", "prompt": "Clean editorial bar chart comparing GLM-4.5-Air output token price via OpenRouter ($0.85/1M) against GLM-4.6 ($1.74/1M) and DeepSeek V4-Flash ($0.28/1M), annotated with the live-measured single-call cost of $0.00017374 for Air. Teal and amber palette, minimal background, 16:9.", "alt": "Bar chart comparing GLM-4.5-Air, GLM-4.6, and DeepSeek V4-Flash output token prices, annotated with the live-measured single-call cost of $0.00017374 for GLM-4.5-Air via OpenRouter"} ] } -->

Share: