Yes, GLM-4.7-Flash is genuinely free on Z.ai International, where input and output both bill at zero, but the same model costs $0.06 / $0.40 per million tokens when you route it through OpenRouter, and free is not the same as fast. That tradeoff is the whole story, so we tested it.
The question buyers actually ask is simpler than the pricing pages make it sound: is the free tier real, or is it the kind of "free" that throttles you into a paid plan by Tuesday? On the official platform the answer is the good kind. On the reseller route it is not free at all. Knowing which door you walked through decides your bill.
According to Z.ai pricing, GLM-4.7-Flash and GLM-4.5-Flash are offered on a free tier on the Z.ai international platform, with zero cost on both input and output tokens. This is a true free tier, not a trial credit that expires, and it is the official-platform fact you should anchor on. We flag it as needing native-key re-verification because we have not yet billed a call against a Z.ai key directly, only confirmed the published rate.
GLM-4.7-Flash is a lightweight model in Zhipu's GLM family that trades raw capability for a price point low enough to run at scale. The "Flash" label signals the budget tier, the same way DeepSeek's V4-Flash sits below V4-Pro.
| Route | Input ($/1M) | Output ($/1M) | Notes |
|---|---|---|---|
| Z.ai International (official) | $0.00 | $0.00 | Free tier, per Z.ai pricing |
| OpenRouter (reseller) | $0.06 | $0.40 | Measured via OpenRouter catalog |
The gap is not a rounding error. One path is free, the other charges real money for the identical model. Resellers do not get access to the vendor's free tier; they buy capacity and mark it up, so a model that costs nothing at the source can still cost you something downstream.
Free or cheap does not mean quick, and our own run made that uncomfortably clear. When we called GLM-4.7-Flash through OpenRouter on 2026-06-26, the request took 26.24 seconds to return 200 output tokens against a 31-token prompt, and OpenRouter billed it at $0.000103875. That was the slowest call in our entire GLM test set.
For context, in the same session GLM-4.6 returned its 200-token cap in 8.58 seconds, GLM-4.5-Air in 8.07 seconds, and GLM-5-Turbo in 7.66 seconds. The "Flash" tier, the one whose name promises speed, came in dead last at more than three times the latency of the next-slowest model. We measured this via OpenRouter, not a native Z.ai key, so the absolute latency may differ on the official endpoint, but the ranking surprised us enough that we noted it.
| Model (via OpenRouter) | Output tokens | Latency | Billed cost |
|---|---|---|---|
| GLM-5-Turbo | 200 | 7.66s | $0.0008372 |
| GLM-4.5-Air | 200 | 8.07s | $0.00017374 |
| GLM-4.6 | 200 | 8.58s | $0.00036376 |
| GLM-4.7-Flash | 200 | 26.24s | $0.000103875 |
The lesson generalizes. We saw the same "Flash naming does not equal fast response" pattern earlier with Qwen's flash tier, which is why we now treat the speed claim in any "Flash" model name as a hypothesis to test, not a spec to trust.
Reach for the free GLM-4.7-Flash tier when your workload is latency-tolerant and cost-sensitive: batch summarization, overnight data enrichment, draft generation, classification jobs where a few extra seconds per call cost you nothing. At zero per token on the official platform, the economics are unbeatable for anything that does not sit in front of a waiting user.
Avoid the free tier, and pay for a faster model, when latency is the product. If a human is watching a cursor blink, a 26-second response is a lost user, and the few cents you save per million tokens will not buy that user back. In that case GLM-4.6 or GLM-5-Turbo, both of which returned in under nine seconds in our test, are the safer spend.
According to Zhipu AI open platform docs, the GLM family spans these tiers precisely so you can match model weight to workload, and the Flash tier exists for high-volume, cost-bound use rather than interactive speed. The practical upshot is that "which GLM model" is really a question about whether your job waits well.
If the free tier is unavailable to you, or you need the request answered now, DeepSeek is the obvious budget comparator. According to DeepSeek API Docs, DeepSeek V4-Flash bills at $0.14 input and $0.28 output per million tokens on its official endpoint. That is not free, but when we called DeepSeek V4-Flash directly on its official API, it returned a 120-token coding answer in 2.25 seconds, more than ten times faster than our GLM-4.7-Flash call.
So the real decision is rarely "GLM-4.7-Flash free versus paid GLM-4.7-Flash." It is "GLM free and slow" versus "DeepSeek cheap and fast." Note the asymmetry in our evidence: the GLM number is OpenRouter-routed while the DeepSeek number is from the official endpoint, so we are comparing a reseller-routed GLM latency against a first-party DeepSeek latency. The gap is wide enough that the disclosure does not erase it, but it is worth stating plainly. For a fuller DeepSeek breakdown, see our DeepSeek API pricing hub.
Is GLM-4.7-Flash really free? On the official Z.ai international platform, yes: GLM-4.7-Flash and GLM-4.5-Flash bill at zero for both input and output per Z.ai pricing. It is a standing free tier, not an expiring trial. But if you reach the model through a reseller such as OpenRouter, you pay their rate, which was $0.06 input and $0.40 output per million tokens when we checked on 2026-06-26.
Why does OpenRouter charge for a free model? Resellers do not inherit the vendor's free tier. They buy inference capacity and resell it with a margin, so the GLM-4.7-Flash that costs nothing on Z.ai's own platform costs $0.06 / $0.40 per million through OpenRouter. The model is identical; only the billing path differs.
Is GLM-4.7-Flash fast? Not in our test. Measured via OpenRouter on 2026-06-26, GLM-4.7-Flash took 26.24 seconds to return 200 tokens, the slowest of the five GLM tiers we called. The "Flash" name describes its budget position, not its response speed.
What should I use if I need speed instead of a free tier? A heavier GLM tier or DeepSeek. GLM-4.6 and GLM-5-Turbo both returned in under nine seconds in our OpenRouter run, and DeepSeek V4-Flash answered in 2.25 seconds on its official endpoint at $0.14 / $0.28 per million tokens.
This is part of the GLM API pricing hub, which compares every GLM tier and the routes that serve them.
Author: Kevin Fan, Customer Success Manager at China LLM Directory, specializing in Chinese LLM ecosystem pricing. Last verified: 2026-06-26.
<!-- METADATA { "title": "GLM-4.7-Flash: Is It Really Free? (2026)", "slug": "glm-4.7-flash-free-tier", "meta_description": "GLM-4.7-Flash is free on Z.ai International (zero in/out) but $0.06/$0.40 via OpenRouter. We tested it: the free tier was also the slowest, at 26.24s. 2026-06.", "focus_keyword": "glm-4.7-flash free tier", "secondary_keywords": ["is glm-4.7-flash free", "glm-4.5-flash free", "glm flash pricing", "z.ai free tier", "glm-4.7-flash openrouter cost"], "tags": ["GLM", "Zhipu", "API Pricing", "Free Tier"], "category": "Pricing", "cluster_id": "glm-api-pricing", "cluster_role": "micro", "hub_slug": "glm-api-pricing", "evidence_file": "clients/china-llm-aggregator/articles/glm-api-pricing-evidence.json", "needs_native_reverify": true, "verified_until": "2026-09-24", "author_name": "Kevin Fan", "author_title": "Customer Success Manager", "author_linkedin": "", "author_expertise": ["Chinese LLM ecosystem", "AI infrastructure pricing", "model benchmarking", "cross-border AI compliance"], "faq_pairs": [ {"q": "Is GLM-4.7-Flash really free?", "a": "On the official Z.ai international platform, yes: GLM-4.7-Flash and GLM-4.5-Flash bill at zero for both input and output per Z.ai pricing. It is a standing free tier, not an expiring trial. But if you reach the model through a reseller such as OpenRouter, you pay their rate, which was $0.06 input and $0.40 output per million tokens when we checked on 2026-06-26."}, {"q": "Why does OpenRouter charge for a free model?", "a": "Resellers do not inherit the vendor's free tier. They buy inference capacity and resell it with a margin, so the GLM-4.7-Flash that costs nothing on Z.ai's own platform costs $0.06 / $0.40 per million through OpenRouter. The model is identical; only the billing path differs."}, {"q": "Is GLM-4.7-Flash fast?", "a": "Not in our test. Measured via OpenRouter on 2026-06-26, GLM-4.7-Flash took 26.24 seconds to return 200 tokens, the slowest of the five GLM tiers we called. The Flash name describes its budget position, not its response speed."}, {"q": "What should I use if I need speed instead of a free tier?", "a": "A heavier GLM tier or DeepSeek. GLM-4.6 and GLM-5-Turbo both returned in under nine seconds in our OpenRouter run, and DeepSeek V4-Flash answered in 2.25 seconds on its official endpoint at $0.14 / $0.28 per million tokens."} ], "external_links_used": [ {"url": "https://docs.z.ai/", "source_name": "Z.ai pricing", "claim": "GLM-4.7-Flash and GLM-4.5-Flash free tier (zero in/out) on Z.ai international"}, {"url": "https://docs.bigmodel.cn/", "source_name": "Zhipu AI open platform docs", "claim": "GLM family tier structure; Flash tier targets high-volume cost-bound use"}, {"url": "https://api-docs.deepseek.com/quick_start/pricing/", "source_name": "DeepSeek API Docs", "claim": "DeepSeek V4-Flash official pricing $0.14 input / $0.28 output per million tokens"} ], "internal_links_used": [ {"url": "/blog/glm-api-pricing/", "anchor_text": "GLM API pricing hub", "type": "hub"}, {"url": "/blog/deepseek-api-pricing/", "anchor_text": "DeepSeek API pricing hub", "type": "cross-cluster"} ], "first_hand_evidence": { "source": "glm-api-pricing-evidence.json run glm-4.7-flash_general + deepseek_v4flash_coding_anchor", "measured": "GLM-4.7-Flash via OpenRouter: 31 in / 200 out, $0.000103875, 26.24s (slowest of 5 GLM tiers); GLM-4.6 8.58s, GLM-4.5-Air 8.07s, GLM-5-Turbo 7.66s; DeepSeek V4-Flash official 27 in / 120 out, 2.25s", "captured": "2026-06-26", "disclosure": "GLM values OpenRouter-routed; DeepSeek value from official api.deepseek.com endpoint" }, "images_status": "spec-only (not generated; FAL_API_KEY unset)", "images": [ {"position": "featured", "type": "generated", "prompt": "Clean editorial diagram contrasting two billing paths for GLM-4.7-Flash: a free Z.ai International door at $0.00 and a paid OpenRouter door at $0.06/$0.40, with a clock annotation showing the measured 26.24s latency. Blue and amber palette. Minimal background. 16:9.", "alt": "Diagram contrasting GLM-4.7-Flash free Z.ai International tier against the paid OpenRouter route, annotated with the measured 26.24-second latency"} ] } -->