GLM-4.6 runs $0.43 in / $1.74 out per million tokens routed through OpenRouter, undercutting Z.ai's own $0.60 / $2.20 official rate, while Z.ai direct offers a genuinely free GLM-4.7-Flash tier OpenRouter does not match. That split is the whole story for international buyers: the cheapest paid GLM lives on a third-party route, and the only free GLM lives on the vendor's own platform. This hub maps both levers and links every deep-dive in the cluster.
Most GLM pricing pages quote one number and stop. The reality buyers hit is two separate markets for the same models. A reseller route shaves 25 to 35 percent off paid tiers, and the vendor's free Flash tier takes a low-volume workload to zero, but you cannot get both from one place. Picking the right tier and route is the decision this page settles.
According to Z.ai Pricing Docs, Zhipu's GLM family spans a free Flash tier up through the flagship reasoning models, and each article below answers one buyer question with verified prices and a first-hand test. GLM is the open-weight model family from Zhipu AI (Z.ai) that competes with DeepSeek and GPT-4o on price-per-token rather than on brand recognition.
| Sub-topic | Article | The number that matters |
|---|---|---|
| Flagship tier rate | GLM-4.6 API pricing | $0.43 in / $1.74 out via OpenRouter vs $0.60 / $2.20 official |
| Cheap mid tier | GLM-4.5-Air pricing | $0.13 in / $0.85 out via OpenRouter |
| The free option | GLM-4.7-Flash free tier | $0 on Z.ai direct; OpenRouter charges $0.06 / $0.40 |
| GLM vs DeepSeek | GLM vs DeepSeek cost | DeepSeek V4-Flash $0.14 / $0.28 official undercuts paid GLM |
| GLM vs GPT-4o | GLM vs GPT-4o cost | GPT-4o $2.50 / $10 vs GLM-4.6 $0.43 / $1.74 |
| Absolute cheapest | Cheapest GLM model | free Flash on Z.ai, then Air at $0.13 / $0.85 |
| Coding workloads | GLM coding cost | GLM-4.6 coding call billed $0.0002972 in 3.78s |
| US access | Is GLM available in the US? | yes, via Z.ai international; free tier included |
| Estimate your bill | GLM API cost calculator | plug your token mix into the per-tier rates |
Sources: Z.ai Pricing Docs and DeepSeek API Docs – Pricing, verified 2026-06-26.
Here is the part most buyers miss. The same GLM weights are sold through two channels at materially different prices, and neither channel is strictly better. The OpenRouter route is cheaper on every paid tier we checked. The Z.ai direct route is the only one with a free tier.
According to Z.ai Pricing Docs, GLM-4.6 lists at $0.60 input and $2.20 output per million tokens on the vendor's own platform. The OpenRouter-routed price for the same model sits at $0.43 / $1.74, roughly 28 percent cheaper on input and 21 percent cheaper on output. We have seen this exact inversion with Qwen too. It means the "official" price is not the floor.
The flip side is the free tier. GLM-4.7-Flash and GLM-4.5-Flash cost nothing on Z.ai international, zero on both input and output, while OpenRouter charges $0.06 / $0.40 for the same 4.7-Flash. So the cheapest paid GLM is on OpenRouter, but the cheapest GLM overall is the free Flash tier on Z.ai direct.
| Model | OpenRouter in / out ($/1M) | Z.ai official in / out ($/1M) | Cheaper route |
|---|---|---|---|
| GLM-4.7-Flash | $0.06 / $0.40 | free | Z.ai (free) |
| GLM-4.5-Air | $0.13 / $0.85 | $0.20 / $1.10 | OpenRouter |
| GLM-4.6 | $0.43 / $1.74 | $0.60 / $2.20 | OpenRouter |
| GLM-5-Turbo | $1.20 / $4.00 | needs reverify | OpenRouter |
OpenRouter figures are the values we measured live on the OpenRouter catalog on 2026-06-26. The Z.ai official column is sourced from Zhipu documentation and carries a needs_native_reverify flag, since we reached GLM through OpenRouter rather than a native Z.ai key. Never read the OpenRouter column as Z.ai's posted price; they are two different markets.
Documentation gives you a rate card. A live call tells you what the rate card does not. We routed five GLM tiers through OpenRouter on 2026-06-26 and billed real tokens against real latency, and one result genuinely surprised us.
The GLM-4.7-Flash general call was the slowest in the entire set. It returned 200 output tokens against a 31-token prompt and took 26.24 seconds, billed at $0.000103875, measured via OpenRouter. By comparison, GLM-4.6 answered a similar prompt in 8.58 seconds and GLM-5-Turbo in 7.66 seconds. We re-read the run to be sure, because a tier named "Flash" being three times slower than the flagship is the opposite of what the name promises. We saw the same "flash naming does not mean fast" pattern with Qwen3.5-Flash, so this is not a one-off.
| GLM tier (via OpenRouter) | Prompt / output tokens | Billed cost | Latency |
|---|---|---|---|
| GLM-4.6 general | 32 / 200 | $0.00036376 | 8.58s |
| GLM-4.6 coding | 30 / 165 | $0.0002972 | 3.78s |
| GLM-4.5-Air general | 32 / 200 | $0.00017374 | 8.07s |
| GLM-4.7-Flash general | 31 / 200 | $0.000103875 | 26.24s |
| GLM-5-Turbo general | 31 / 200 | $0.0008372 | 7.66s |
The practical upshot: do not pick the Flash tier expecting speed. On Z.ai its appeal is the zero price, not the latency. If response time is your constraint, GLM-4.6 or GLM-5-Turbo served us far faster. All GLM general calls hit the 200-token output cap, so treat that length as capped. Every number above is OpenRouter-routed and live-billed.
No single sub-article answers the question buyers actually arrive with, which is "given my situation, which GLM and which route do I pick?" Here is the cross-cutting view that ties the cluster together.
| Your situation | Tier | Route | Why |
|---|---|---|---|
| Prototype, hobby, or bursty low volume | GLM-4.7-Flash | Z.ai direct (free) | zero cost beats any paid rate; accept the slow latency |
| Production chat, cost-sensitive | GLM-4.5-Air | OpenRouter | $0.13 / $0.85 is the cheapest competent paid tier |
| Quality-sensitive flagship work | GLM-4.6 | OpenRouter | $0.43 / $1.74 undercuts Z.ai's own $0.60 / $2.20 |
| Heaviest reasoning, latency matters | GLM-5-Turbo | OpenRouter | fastest flagship we measured at 7.66s |
| Regulated or US-residency requirement | depends | Z.ai international or US reseller | route is a compliance gate, see the US-access article |
The ordering matters. Start with the free tier question: if your volume is low enough that GLM-4.7-Flash free on Z.ai covers it, nothing on OpenRouter beats zero. Only when free-tier limits or the 26-second latency become real constraints does the paid decision begin, and there OpenRouter wins on price for every tier we tested. Route choice is the last lever unless compliance forces it earlier.
The cross-model picture changes the verdict. Blended cost is the combined input-plus-output charge for your actual token mix, and it is the figure to budget against rather than any single rate. According to DeepSeek API Docs – Pricing, DeepSeek V4-Flash runs $0.14 / $0.28 on its official endpoint, which undercuts paid GLM-4.6 even at the discounted OpenRouter rate. So GLM's price edge is real against GPT-4o at $2.50 / $10, but it is not the cheapest Chinese model for paid work; DeepSeek holds that spot. GLM's distinct lever is the free Flash tier. We anchored this with a live DeepSeek V4-Flash coding call that returned in 2.25 seconds on the official endpoint, far faster than the GLM-via-OpenRouter calls.
According to Ropes & Gray LLP, enterprise buyers should treat endpoint routing as a compliance decision separate from price, which is why our matrix keeps the US-residency row independent of tier.
Prices here were pulled from Z.ai and DeepSeek documentation and confirmed against live API calls on 2026-06-26. GLM tiers were reached through OpenRouter, not a native Z.ai key, so every GLM measurement is OpenRouter-routed and the official Z.ai figures carry a needs_native_reverify flag. The aggregator does not host or resell any model API; we publish neutral comparisons and refresh prices when a provider's documentation changes.
How much does the GLM-4.6 API cost? GLM-4.6 costs $0.43 per million input tokens and $1.74 per million output tokens when routed through OpenRouter, which we measured live on 2026-06-26. Z.ai's own official rate is higher at $0.60 / $2.20, so the marketplace route is roughly 25 to 28 percent cheaper for this flagship tier.
Is there a free GLM API tier? Yes. GLM-4.7-Flash and GLM-4.5-Flash are free on the Z.ai international platform, zero cost on both input and output. This is a genuine vendor free tier that OpenRouter does not match; OpenRouter charges $0.06 / $0.40 for the same 4.7-Flash model.
Is GLM cheaper than GPT-4o? Substantially. GPT-4o lists at $2.50 input and $10 output per million tokens, while GLM-4.6 via OpenRouter is $0.43 / $1.74. That is roughly a 6x saving on input and a 5.7x saving on output, before considering GLM's free Flash tier for low-volume work.
Is GLM cheaper than DeepSeek? Not for paid work. DeepSeek V4-Flash runs $0.14 / $0.28 on its official endpoint, which undercuts even the discounted GLM-4.6 OpenRouter rate. GLM's advantage over DeepSeek is the free Flash tier, not a lower paid price.
Why is the GLM Flash tier so slow? In our live test, GLM-4.7-Flash took 26.24 seconds to return 200 tokens, the slowest of the five tiers we measured via OpenRouter, while GLM-4.6 answered in 8.58 seconds. The "Flash" name signals the zero price on Z.ai, not low latency, so choose it for cost rather than speed.
Can I use GLM from the US? Yes. GLM is reachable from the US through the Z.ai international platform, and the free Flash tier is available to international developers. For regulated workloads, treat the endpoint as a compliance decision; see the US-access article in this cluster for the residency detail.
Should I use OpenRouter or Z.ai direct for GLM? Use Z.ai direct if the free Flash tier covers your volume, since nothing beats zero. Use OpenRouter for any paid tier, because it undercut Z.ai's official price on every model we checked. The two routes serve different needs and you generally cannot get both advantages at once.
Author: Kevin Fan, Customer Success Manager at China LLM Directory, specializing in Chinese LLM ecosystem pricing. Last verified: 2026-06-26.
<!-- METADATA { "title": "GLM (Zhipu) API Pricing: Buyer Guide 2026", "slug": "glm-api-pricing", "meta_description": "GLM-4.6 is $0.43/$1.74 via OpenRouter vs $0.60/$2.20 official Z.ai, plus a free GLM-4.7-Flash tier. Verified rates, live tests, route matrix. 2026-06.", "focus_keyword": "glm api pricing", "secondary_keywords": ["glm 4.6 pricing", "zhipu api pricing", "glm free tier", "glm vs deepseek cost", "z.ai pricing"], "tags": ["GLM", "Zhipu", "API Pricing"], "category": "Pricing", "cluster_id": "glm-api-pricing", "cluster_role": "hub", "evidence_file": "clients/china-llm-aggregator/articles/glm-api-pricing-evidence.json", "needs_native_reverify": true, "verified_until": "2026-09-24", "author_name": "Kevin Fan", "author_title": "Customer Success Manager", "author_linkedin": "", "author_expertise": ["Chinese LLM ecosystem", "AI infrastructure pricing", "model benchmarking", "cross-border AI compliance"], "faq_pairs": [ {"q": "How much does the GLM-4.6 API cost?", "a": "GLM-4.6 costs $0.43 per million input tokens and $1.74 per million output tokens when routed through OpenRouter, which we measured live on 2026-06-26. Z.ai's own official rate is higher at $0.60 / $2.20, so the marketplace route is roughly 25 to 28 percent cheaper for this flagship tier."}, {"q": "Is there a free GLM API tier?", "a": "Yes. GLM-4.7-Flash and GLM-4.5-Flash are free on the Z.ai international platform, zero cost on both input and output. This is a genuine vendor free tier that OpenRouter does not match; OpenRouter charges $0.06 / $0.40 for the same 4.7-Flash model."}, {"q": "Is GLM cheaper than GPT-4o?", "a": "Substantially. GPT-4o lists at $2.50 input and $10 output per million tokens, while GLM-4.6 via OpenRouter is $0.43 / $1.74. That is roughly a 6x saving on input and a 5.7x saving on output, before considering GLM's free Flash tier for low-volume work."}, {"q": "Is GLM cheaper than DeepSeek?", "a": "Not for paid work. DeepSeek V4-Flash runs $0.14 / $0.28 on its official endpoint, which undercuts even the discounted GLM-4.6 OpenRouter rate. GLM's advantage over DeepSeek is the free Flash tier, not a lower paid price."}, {"q": "Why is the GLM Flash tier so slow?", "a": "In our live test, GLM-4.7-Flash took 26.24 seconds to return 200 tokens, the slowest of the five tiers we measured via OpenRouter, while GLM-4.6 answered in 8.58 seconds. The Flash name signals the zero price on Z.ai, not low latency, so choose it for cost rather than speed."}, {"q": "Can I use GLM from the US?", "a": "Yes. GLM is reachable from the US through the Z.ai international platform, and the free Flash tier is available to international developers. For regulated workloads, treat the endpoint as a compliance decision; see the US-access article in this cluster for the residency detail."}, {"q": "Should I use OpenRouter or Z.ai direct for GLM?", "a": "Use Z.ai direct if the free Flash tier covers your volume, since nothing beats zero. Use OpenRouter for any paid tier, because it undercut Z.ai's official price on every model we checked. The two routes serve different needs and you generally cannot get both advantages at once."} ], "external_links_used": [ {"url": "https://docs.z.ai/guides/overview/pricing", "source_name": "Z.ai Pricing Docs", "claim": "GLM-4.6 official $0.60/$2.20; GLM-4.7-Flash free tier; full GLM rate card"}, {"url": "https://api-docs.deepseek.com/quick_start/pricing/", "source_name": "DeepSeek API Docs – Pricing", "claim": "DeepSeek V4-Flash $0.14/$0.28 official endpoint, comparator anchor"}, {"url": "https://www.ropesgray.com/en/insights/alerts/2025/01/deepseek-legal-considerations-for-enterprise-users", "source_name": "Ropes & Gray LLP", "claim": "Enterprise endpoint routing is a compliance decision separate from price"} ], "internal_links_used": [ {"url": "/blog/glm-4.6-pricing/", "anchor_text": "GLM-4.6 API pricing", "type": "micro"}, {"url": "/blog/glm-4.5-air-pricing/", "anchor_text": "GLM-4.5-Air pricing", "type": "micro"}, {"url": "/blog/glm-4.7-flash-free-tier/", "anchor_text": "GLM-4.7-Flash free tier", "type": "micro"}, {"url": "/blog/glm-vs-deepseek-cost/", "anchor_text": "GLM vs DeepSeek cost", "type": "micro"}, {"url": "/blog/glm-vs-gpt-4o-cost/", "anchor_text": "GLM vs GPT-4o cost", "type": "micro"}, {"url": "/blog/cheapest-glm-model/", "anchor_text": "Cheapest GLM model", "type": "micro"}, {"url": "/blog/glm-coding-cost/", "anchor_text": "GLM coding cost", "type": "micro"}, {"url": "/blog/is-glm-available-in-us/", "anchor_text": "Is GLM available in the US?", "type": "micro"}, {"url": "/blog/glm-api-cost-calculator/", "anchor_text": "GLM API cost calculator", "type": "micro"} ], "first_hand_evidence": { "source": "glm-api-pricing-evidence.json runs glm-4.7-flash_general + glm-4.6_general + deepseek_v4flash_coding_anchor", "measured": "GLM-4.7-Flash via OpenRouter 31/200 tokens, $0.000103875, 26.24s (slowest of set); GLM-4.6 8.58s; GLM-5-Turbo 7.66s; DeepSeek V4-Flash coding anchor 2.25s official endpoint", "captured": "2026-06-26", "routing": "GLM tiers OpenRouter-routed; DeepSeek anchor official api.deepseek.com" }, "images_status": "spec-only (not generated; FAL_API_KEY unset)", "images": [ {"position": "featured", "type": "generated", "prompt": "Clean editorial diagram of the GLM model tier ladder (4.7-Flash free, 4.5-Air, 4.6, 5-Turbo) with two pricing routes branching off each tier: OpenRouter paid and Z.ai direct (free for Flash). Annotated with the live datapoint that GLM-4.7-Flash measured 26 seconds. Blue and amber palette. Minimal background. 16:9.", "alt": "Diagram of GLM model tiers split across two pricing routes, OpenRouter paid and Z.ai direct free Flash tier, annotated with the 26-second GLM-4.7-Flash latency datapoint"} ] } -->