Is GLM Available in the US? (2026)

Yes, GLM from Zhipu AI is available in the US through the Z.ai international platform, which offers a free Flash tier, and you can also reach GLM models…

Fan Chuanyu's profile

Written by Fan Chuanyu

8 min read

Yes, GLM from Zhipu AI is available in the US through the Z.ai international platform, which offers a free Flash tier, and you can also reach GLM models through OpenRouter routing if you prefer a multi-provider gateway. There is no US ban on signing up; access is a billing-and-compliance question, not a technical one.

If you have read our DeepSeek US access article, the shape of this answer will feel familiar. A China-based lab ships a strong, cheap model; it is reachable from the US; the real decision is not "can I" but "should my organization route data to a China-headquartered provider." GLM is the same pattern with its own specifics.

How to access GLM (Zhipu) from the US (verified 2026-06)

GLM is the model family from Zhipu AI, a Beijing-based research lab whose international developer product is branded Z.ai. There are two practical front doors from a US location.

The first is the Z.ai international platform directly. You create an account, get a native API key, and call GLM-4.6, the Air variant, and the Flash tier. According to Z.ai developer documentation, the international platform is the intended path for developers outside mainland China, and it exposes the same GLM model lineup through an OpenAI-compatible endpoint.

The second is a multi-provider gateway. We reached every GLM tier in this article through OpenRouter rather than a native Z.ai key, which is the honest disclosure behind our first-hand numbers below. Gateways are a reseller layer, not the source, so the price you pay there is set by the router and not by Zhipu.

GLM is a large language model family that competes with DeepSeek and Qwen on the price-to-capability frontier, which is the differentiator most US buyers care about: it lands well below US-hosted frontier models on cost while staying close on coding and general reasoning.

What it costs: official Z.ai versus gateway routing

Here is the part that trips people up. The price you see depends entirely on which door you walked through, and the two do not match.

GLM tierZ.ai official (in / out per 1M)Measured via OpenRouter (in / out per 1M)
GLM-4.6$0.60 / $2.20$0.43 / $1.74
GLM-4.5-Air$0.20 / $1.10$0.13 / $0.85
GLM-4.7-FlashFREE (Z.ai international free tier)$0.06 / $0.40

The official rates are sourced research that still needs native-key re-verification, so treat them as directional. According to Z.ai pricing, GLM-4.6 is listed at $0.60 input and $2.20 output per million tokens on a 205K-context tier. The gateway-routed price we observed for the same model was lower, $0.43 in and $1.74 out, which is the same pattern we documented when a router undercuts the lab's own list price. That gap is the part most buyers miss: cheaper on a gateway does not mean cheaper from the source, it means the router chose to discount.

The headline for US access, though, is the free tier. According to Z.ai pricing, GLM-4.7-Flash and GLM-4.5-Flash run at zero cost on the international platform. That is a genuine free tier from the lab itself, and a gateway will not give it to you; routed through OpenRouter, the same 4.7-Flash model billed us $0.06 in and $0.40 out. If your goal is to evaluate GLM from a US desk without putting a card down, the native Z.ai Flash tier is the answer, not the gateway.

First-hand evidence: what GLM actually cost us (measured via OpenRouter)

Documentation tells you the rate card. We wanted to see real tokens, real latency, and a real billed cost, so we called the models ourselves. Every GLM number in this section was measured via OpenRouter on 2026-06-26, not through a native Z.ai key, and we disclose that because it changes how you should read the cost.

When we called GLM-4.6 with a short general prompt, the request used 32 input tokens and returned 200 output tokens, billed at $0.00036376, with a latency of 8.58 seconds. The Air variant on the same general prompt came in cheaper at $0.00017374 for 32 in and 200 out, taking 8.07 seconds.

The surprise was the Flash tier. You expect a model named "Flash" to be the quick one. Routed through OpenRouter, GLM-4.7-Flash on a 31-in, 200-out general call billed only $0.000103875, the cheapest call in the set, but it took 26.24 seconds, by far the slowest response we measured. We re-read the run to be sure. The lesson, and we saw the identical thing with a Qwen flash tier, is that "flash" naming describes the price class, not the wall-clock speed you will get on a given route.

GLM model (via OpenRouter)Tokens (in / out)Billed cost (USD)Latency
GLM-4.6 (general)32 / 200$0.000363768.58s
GLM-4.5-Air (general)32 / 200$0.000173748.07s
GLM-4.7-Flash (general)31 / 200$0.00010387526.24s

One caveat we will not hide: all of these general calls hit a 200-token output cap, so treat the 200 completion tokens as capped, not a model's natural answer length. The costs are real billed amounts; the output length is a harness artifact.

How GLM access compares to DeepSeek for a US buyer

For a US team weighing China-based options, GLM and DeepSeek sit in the same risk bucket but with different cost mechanics. According to DeepSeek API documentation, DeepSeek V4-Flash bills $0.14 input and $0.28 output per million on its official endpoint, which is reached natively rather than through a router. Our DeepSeek anchor call, made against that official endpoint, returned 120 output tokens for a 27-token prompt in 2.25 seconds, a notably faster turnaround than our gateway-routed GLM calls.

So the asymmetry is worth naming: our DeepSeek figure is official-endpoint, while our GLM figures are gateway-routed. Choose GLM if you want a free Flash tier to prototype on. Choose DeepSeek if you want a low, predictable official rate and a single native billing relationship. Both leave you with the same data-residency question.

Compliance and data residency: the question that actually matters

Zhipu AI is headquartered in China, and that fact drives the procurement conversation more than any price. Inputs and outputs sent to the Z.ai international platform are processed by a China-based provider, which for regulated buyers in finance, healthcare, or government work usually triggers a data-residency and vendor-risk review before any production traffic flows. The free Flash tier lowers the cost of evaluation to zero, but it does not change the residency posture; a free request still leaves your environment.

The practical upshot for most US teams is a two-stage approach. Prototype on the free Flash tier to judge whether GLM's quality fits your task, treating prompts as non-sensitive during that phase. Then, before sending real customer or proprietary data, run the vendor review you would apply to any offshore processor. Access is the easy part. The review is the part that protects you.

FAQ

Is GLM (Zhipu) banned in the US? No. There is no US ban on accessing GLM. You can sign up for the Z.ai international platform from a US location and call the models, and you can also reach them through a multi-provider gateway. The constraint is your own organization's data-residency and vendor-risk policy, not US law.

Does GLM have a free tier I can use from the US? Yes. According to Z.ai pricing, GLM-4.7-Flash and GLM-4.5-Flash run at zero cost on the Z.ai international platform. This is a genuine lab-provided free tier. Note this is an official Z.ai fact and is pending native-key re-verification; routed through a gateway instead, the same Flash model is billed, not free.

How much does GLM-4.6 cost? It depends on the door. According to Z.ai pricing, GLM-4.6 is listed at $0.60 input and $2.20 output per million tokens. Measured via OpenRouter, we saw $0.43 in and $1.74 out for the same model, lower because the router sets its own price. Always confirm which path a quoted number came from.

Is sending data to GLM a compliance risk for a US company? It can be. Zhipu AI is a China-based provider, so data sent to Z.ai is processed outside the US. Regulated buyers should run a data-residency and vendor-risk review before production use, exactly as you would with our parallel guidance on DeepSeek.

Is GLM faster than DeepSeek? In our tests, no. Our DeepSeek V4-Flash anchor call on the official endpoint returned in 2.25 seconds, while our gateway-routed GLM calls ranged from 3.78 to 26.24 seconds. Routing differences explain much of that gap, so treat latency as path-dependent, not as a fixed model property.


This is part of the GLM API pricing hub, where we collect every verified GLM rate and benchmark in one place.

Author: Kevin Fan, Customer Success Manager at China LLM Directory, specializing in Chinese LLM ecosystem pricing and cross-border AI compliance. Last verified: 2026-06-26.

<!-- METADATA { "title": "Is GLM (Zhipu) Available in the US? (2026)", "slug": "is-glm-available-in-us", "meta_description": "Yes, GLM from Zhipu is available in the US via the Z.ai international platform with a free Flash tier, plus OpenRouter routing. Access, pricing, and compliance.", "focus_keyword": "is glm available in us", "secondary_keywords": ["glm zhipu us access", "z.ai international platform", "glm free tier us", "glm api compliance us"], "tags": ["GLM", "Zhipu", "US Access", "API Pricing", "Compliance"], "category": "Access", "cluster_id": "glm-api-pricing", "cluster_role": "micro", "hub_slug": "glm-api-pricing", "evidence_file": "clients/china-llm-aggregator/articles/glm-api-pricing-evidence.json", "needs_native_reverify": true, "verified_until": "2026-09-24", "author_name": "Kevin Fan", "author_title": "Customer Success Manager", "author_linkedin": "", "author_expertise": ["Chinese LLM ecosystem", "AI infrastructure pricing", "model benchmarking", "cross-border AI compliance"], "faq_pairs": [ {"q": "Is GLM (Zhipu) banned in the US?", "a": "No. There is no US ban on accessing GLM. You can sign up for the Z.ai international platform from a US location and call the models, and you can also reach them through a multi-provider gateway. The constraint is your own organization's data-residency and vendor-risk policy, not US law."}, {"q": "Does GLM have a free tier I can use from the US?", "a": "Yes. According to Z.ai pricing, GLM-4.7-Flash and GLM-4.5-Flash run at zero cost on the Z.ai international platform. This is a genuine lab-provided free tier, pending native-key re-verification. Routed through a gateway instead, the same Flash model is billed, not free."}, {"q": "How much does GLM-4.6 cost?", "a": "It depends on the door. According to Z.ai pricing, GLM-4.6 is listed at $0.60 input and $2.20 output per million tokens. Measured via OpenRouter, we saw $0.43 in and $1.74 out for the same model, lower because the router sets its own price."}, {"q": "Is sending data to GLM a compliance risk for a US company?", "a": "It can be. Zhipu AI is a China-based provider, so data sent to Z.ai is processed outside the US. Regulated buyers should run a data-residency and vendor-risk review before production use, as with DeepSeek."}, {"q": "Is GLM faster than DeepSeek?", "a": "In our tests, no. Our DeepSeek V4-Flash anchor call on the official endpoint returned in 2.25 seconds, while our gateway-routed GLM calls ranged from 3.78 to 26.24 seconds. Treat latency as path-dependent, not a fixed model property."} ], "external_links_used": [ {"url": "https://docs.z.ai/", "source_name": "Z.ai developer documentation", "claim": "Z.ai international platform is the intended developer path outside mainland China; GLM lineup pricing and free Flash tier"}, {"url": "https://api-docs.deepseek.com/quick_start/pricing/", "source_name": "DeepSeek API Docs - Pricing", "claim": "DeepSeek V4-Flash official rate $0.14 input / $0.28 output per million"} ], "internal_links_used": [ {"url": "/blog/glm-api-pricing/", "anchor_text": "GLM API pricing hub", "type": "hub"}, {"url": "/blog/deepseek-api-pricing/", "anchor_text": "DeepSeek US access article", "type": "cross-cluster"} ], "first_hand_evidence": { "source": "glm-api-pricing-evidence.json runs glm-4.6_general, glm-4.5-air_general, glm-4.7-flash_general, deepseek_v4flash_coding_anchor", "measured": "GLM-4.6 general 32in/200out $0.00036376 8.58s; GLM-4.5-Air general 32in/200out $0.00017374 8.07s; GLM-4.7-Flash general 31in/200out $0.000103875 26.24s (slowest); DeepSeek V4-Flash official anchor 27in/120out 2.25s", "routing": "GLM measured via OpenRouter; DeepSeek anchor via official api.deepseek.com", "captured": "2026-06-26" }, "images_status": "spec-only (not generated; FAL_API_KEY unset)", "images": [ {"position": "featured", "type": "generated", "prompt": "Clean editorial diagram of a US developer accessing GLM (Zhipu) via two paths: the Z.ai international platform with a free Flash tier, and a multi-provider gateway. Annotate the data-residency review step before production. Teal and slate palette. Minimal background. 16:9.", "alt": "Diagram showing a US developer accessing GLM from Zhipu through the Z.ai international platform free Flash tier and a gateway route, with a data-residency review step before production"} ] } -->

Share: