GLM-4.6 routes cheapest through OpenRouter at $0.43 per million input tokens and $1.74 per million output, undercutting Zhipu's own official Z.ai rate of $0.60/$2.20 for the same flagship tier, an unusual case where the third-party aggregator beats the model maker's list price. That gap is the part most buyers miss when they assume going direct is always cheaper. Below we lay out both price points, show a live call we ran, and explain when the official endpoint still wins despite costing more on paper.
GLM-4.6 is Zhipu AI's flagship general-purpose model that pairs a 200K-class context window with frontier-tier reasoning, and it is the tier most teams reach for when GLM-4.5-Air feels too light. Two prices exist for it, and they do not agree, so it is worth being precise about which one you are quoting.
The OpenRouter-routed rate we verified on the live catalog on 2026-06-26 is $0.43 input and $1.74 output per million tokens. Zhipu's own published rate for the same model is higher. The practical upshot: for GLM-4.6 specifically, routing through OpenRouter is the cheaper path, not a convenience tax.
| GLM-4.6 source | Input ($/1M) | Output ($/1M) | Context | Notes |
|---|---|---|---|---|
| OpenRouter-routed | $0.43 | $1.74 | 200K class | Cheaper of the two; what we measured |
| Official Z.ai | $0.60 | $2.20 | 205K | Per Z.ai pricing; needs native re-verify |
According to Z.ai pricing, the official GLM-4.6 rate is $0.60 input and $2.20 output per million tokens with a 205K context window. We carry this as sourced research rather than a first-hand number, because GLM in our test rig is reached through OpenRouter, not a native Zhipu key. That is why this article is flagged for native re-verification.
Documentation gives you a rate card. It does not tell you what a real request costs or how long it takes, so we called the model and watched the meter. Every number here was measured via OpenRouter, billed against the routed $0.43/$1.74 rate, not the official Z.ai card.
On a general-purpose prompt, GLM-4.6 read 32 input tokens, returned 200 output tokens, billed $0.00036376, and took 8.58 seconds. We then sent a coding prompt: 30 input tokens, 165 output tokens, $0.0002972, and a much quicker 3.78 seconds. The coding call finished in under half the time of the general one, which tracks with shorter, more structured completions returning faster.
| GLM-4.6 call (via OpenRouter) | Input tok | Output tok | Billed cost | Latency |
|---|---|---|---|---|
| General | 32 | 200 | $0.00036376 | 8.58s |
| Coding | 30 | 165 | $0.0002972 | 3.78s |
One caveat we always disclose: the general call hit our 200-token output cap, so treat that completion length as capped, not as the model's natural stopping point. The dollar figures are exact billed costs, but a request that runs to a longer natural answer will cost proportionally more on output. The reason we surface our own measurements at all is simple: a token count and a billed cost from a real request are facts a docs-scraping competitor cannot reproduce.
Flagship pricing only means something next to the field. Here is GLM-4.6 against the cheaper GLM tiers and the two non-Chinese frontier models buyers usually weigh it against.
| Model | Input ($/1M) | Output ($/1M) | Source basis |
|---|---|---|---|
| GLM-4.6 | $0.43 | $1.74 | OpenRouter-routed |
| GLM-4.5-Air | $0.13 | $0.85 | OpenRouter-routed |
| GLM-4.7-Flash | FREE | FREE | Official Z.ai free tier |
| DeepSeek V4-Flash | $0.14 | $0.28 | Official api.deepseek.com |
| GPT-4o | $2.50 | $10.00 | OpenAI list |
| Claude Sonnet 4.6 | $3.00 | $15.00 | Anthropic list |
The headline takeaway: GLM-4.6 lands far below GPT-4o and Claude Sonnet 4.6 on both sides of the ledger, roughly 6x cheaper than GPT-4o on input and about 5.7x cheaper on output. Against DeepSeek it is the reverse story. According to DeepSeek API Docs, DeepSeek V4-Flash bills $0.14 input and $0.28 output per million tokens, which makes it markedly cheaper than GLM-4.6 on raw rate, especially on output where DeepSeek is over 6x lower. So GLM-4.6 is not the budget option in the Chinese-model field; it is the mid-priced flagship you pick for capability, with DeepSeek as the value play and the GLM Flash tier as the free experiment.
There is also a genuinely free path inside the GLM family that no aggregator markup touches. According to Z.ai pricing, GLM-4.7-Flash and GLM-4.5-Flash run on a free tier on the Z.ai international platform, billing zero for both input and output. That is an official-Zhipu offering, not an OpenRouter price; OpenRouter itself charges for its 4.7-flash route. If your workload tolerates a lighter model, the Flash free tier is the cheapest entry point in the entire lineup, and GLM-4.6 is the paid upgrade you graduate to when Flash stops being enough.
Budget against the OpenRouter rate of $0.43/$1.74 if you are routing through an aggregator, because that is what you will be billed and it is the cheaper of the two for this tier. Budget against the official $0.60/$2.20 only if you hold a native Z.ai key and want vendor-direct billing, SLA, and data terms.
The decision is not purely about the per-token gap, which is small in absolute dollars. Going direct to Z.ai buys you a single contractual relationship, US and international access through the Z.ai international platform, and the free Flash tier in the same account. Routing through OpenRouter buys you the lower headline rate and one bill across many models. Choose OpenRouter if cost-per-token and multi-model flexibility lead your decision. Choose native Z.ai if you need the free Flash tier, vendor SLA, or want to avoid an intermediary in your data path.
For the full cross-tier breakdown, including GLM-4.5-Air, the Flash free tier, and GLM-5-Turbo, see the GLM API pricing hub. If you are weighing GLM against DeepSeek on cost, our DeepSeek API pricing coverage has the matching first-hand numbers.
How much does the GLM-4.6 API cost? Two rates exist. The OpenRouter-routed rate we verified on 2026-06-26 is $0.43 per million input tokens and $1.74 per million output. According to Z.ai pricing, the official rate is higher at $0.60 input and $2.20 output. For this tier the aggregator route is the cheaper one.
Is GLM-4.6 cheaper on OpenRouter or directly from Z.ai? On OpenRouter, for this specific model. The routed rate of $0.43/$1.74 undercuts the official Z.ai rate of $0.60/$2.20. The official endpoint still wins if you need the free Flash tier, a vendor SLA, or direct billing with no intermediary.
What did a real GLM-4.6 call cost in your test? Measured via OpenRouter, a general-purpose call of 32 input and 200 output tokens billed $0.00036376 and took 8.58 seconds. A coding call of 30 input and 165 output tokens billed $0.0002972 in 3.78 seconds.
How does GLM-4.6 compare to DeepSeek and GPT-4o on price? GLM-4.6 is far below GPT-4o ($2.50/$10) and Claude Sonnet 4.6 ($3/$15), but more expensive than DeepSeek V4-Flash, which bills $0.14/$0.28 per the DeepSeek docs. GLM-4.6 is the mid-priced flagship, not the budget pick.
Is there a free GLM tier? Yes. According to Z.ai pricing, GLM-4.7-Flash and GLM-4.5-Flash run free on the Z.ai international platform for both input and output. It is an official Zhipu offering; OpenRouter charges for its Flash route.
This is part of the GLM API pricing hub, where the full cluster compares every GLM tier side by side.
Author: Kevin Fan, Customer Success Manager at China LLM Directory, specializing in Chinese LLM ecosystem pricing. Last verified: 2026-06-26.
<!-- METADATA { "title": "GLM-4.6 API Pricing Explained: Rates and Costs (2026)", "slug": "glm-4.6-pricing", "meta_description": "GLM-4.6 routes cheapest via OpenRouter at $0.43/$1.74 vs official Z.ai $0.60/$2.20. We measured a live call at $0.00036376 in 8.58s. June 2026.", "focus_keyword": "glm-4.6 api pricing", "secondary_keywords": ["glm-4.6 cost", "glm 4.6 price per token", "glm-4.6 vs deepseek cost", "zhipu glm-4.6 pricing", "glm-4.6 openrouter price"], "tags": ["GLM", "Zhipu", "API Pricing", "Z.ai"], "category": "Pricing", "cluster_id": "glm-api-pricing", "cluster_role": "micro", "hub_slug": "glm-api-pricing", "evidence_file": "clients/china-llm-aggregator/articles/glm-api-pricing-evidence.json", "needs_native_reverify": true, "verified_until": "2026-09-24", "author_name": "Kevin Fan", "author_title": "Customer Success Manager", "author_linkedin": "", "author_expertise": ["Chinese LLM ecosystem", "AI infrastructure pricing", "model benchmarking", "cross-border AI compliance"], "faq_pairs": [ {"q": "How much does the GLM-4.6 API cost?", "a": "Two rates exist. The OpenRouter-routed rate we verified on 2026-06-26 is $0.43 per million input tokens and $1.74 per million output. According to Z.ai pricing, the official rate is higher at $0.60 input and $2.20 output. For this tier the aggregator route is the cheaper one."}, {"q": "Is GLM-4.6 cheaper on OpenRouter or directly from Z.ai?", "a": "On OpenRouter, for this specific model. The routed rate of $0.43/$1.74 undercuts the official Z.ai rate of $0.60/$2.20. The official endpoint still wins if you need the free Flash tier, a vendor SLA, or direct billing with no intermediary."}, {"q": "What did a real GLM-4.6 call cost in your test?", "a": "Measured via OpenRouter, a general-purpose call of 32 input and 200 output tokens billed $0.00036376 and took 8.58 seconds. A coding call of 30 input and 165 output tokens billed $0.0002972 in 3.78 seconds."}, {"q": "How does GLM-4.6 compare to DeepSeek and GPT-4o on price?", "a": "GLM-4.6 is far below GPT-4o ($2.50/$10) and Claude Sonnet 4.6 ($3/$15), but more expensive than DeepSeek V4-Flash, which bills $0.14/$0.28 per the DeepSeek docs. GLM-4.6 is the mid-priced flagship, not the budget pick."}, {"q": "Is there a free GLM tier?", "a": "Yes. According to Z.ai pricing, GLM-4.7-Flash and GLM-4.5-Flash run free on the Z.ai international platform for both input and output. It is an official Zhipu offering; OpenRouter charges for its Flash route."} ], "external_links_used": [ {"url": "https://z.ai/", "source_name": "Z.ai pricing", "claim": "Official GLM-4.6 $0.60/$2.20 with 205K context; GLM-4.7-Flash and GLM-4.5-Flash free tier on Z.ai international"}, {"url": "https://api-docs.deepseek.com/quick_start/pricing/", "source_name": "DeepSeek API Docs", "claim": "DeepSeek V4-Flash bills $0.14 input and $0.28 output per million tokens"} ], "internal_links_used": [ {"url": "/blog/glm-api-pricing/", "anchor_text": "GLM API pricing hub", "type": "hub"}, {"url": "/blog/deepseek-api-pricing/", "anchor_text": "DeepSeek API pricing", "type": "cross-cluster"} ], "first_hand_evidence": { "source": "glm-api-pricing-evidence.json runs glm-4.6_general + glm-4.6_coding", "measured": "glm-4.6 general 32 in / 200 out, $0.00036376, 8.58s; glm-4.6 coding 30 in / 165 out, $0.0002972, 3.78s; all measured via OpenRouter (z-ai/glm-4.6), billed against routed $0.43/$1.74", "captured": "2026-06-26" }, "images_status": "spec-only (not generated; FAL_API_KEY unset)", "images": [ {"position": "featured", "type": "generated", "prompt": "Clean editorial comparison diagram showing GLM-4.6 priced two ways: OpenRouter-routed $0.43/$1.74 versus official Z.ai $0.60/$2.20, with an arrow highlighting the aggregator undercutting the model maker. Teal and slate palette. Minimal background. 16:9.", "alt": "Diagram comparing GLM-4.6 OpenRouter-routed price of $0.43/$1.74 against the official Z.ai price of $0.60/$2.20, showing the aggregator route is cheaper"} ] } -->