Among the major Chinese LLM families, only GLM ships a genuinely free, no-expiry API tier (GLM-4.7-Flash and 4.5-Flash on z.ai); Qwen offers a time-boxed trial of roughly 1M free tokens over 90 days on its Singapore endpoint, while DeepSeek gives new accounts a sign-up credit whose current amount we could not independently verify. That distinction, free forever versus a trial that ends, is the part most "free Chinese LLM" lists blur, so this piece separates the three honestly.
The question behind the keyword is simple: can I actually call a capable Chinese model without paying, and for how long? The answer depends entirely on which family you pick, because "free" means three different things across GLM, Qwen, and DeepSeek.
A free API tier is a usage allowance that recurs without a paywall, distinct from a trial credit that expires once consumed or after a fixed window. Holding that definition straight is what lets us sort the families honestly.
| Family | What is actually free | Type | Expiry | Source basis |
|---|---|---|---|---|
| GLM (z.ai) | GLM-4.7-Flash and GLM-4.5-Flash model calls | Genuinely free tier | None stated | Vendor (z.ai); needs native reverify |
| Qwen (Singapore) | ~1M tokens | Trial | ~90 days | Vendor (Alibaba Cloud Intl); needs native reverify |
| DeepSeek | New-account sign-up credit | Trial credit | Amount unverified | Vendor (api.deepseek.com); unverified |
| Kimi (Moonshot) | No standing free tier we could confirm | Paid-only | n/a | platform.moonshot.ai |
| MiniMax | No standing free tier we could confirm | Paid-only | n/a | platform.minimax.io |
The practical upshot: if you want to keep calling for free past your first month, GLM is the only family on this list that lets you, and even then only on its Flash tiers.
According to z.ai, GLM-4.7-Flash and GLM-4.5-Flash are offered free of charge on the vendor's own endpoint, which makes GLM the only family here with a free tier that does not expire. That matters because a free tier you can build a hobby project on, and a 90-day trial you have to migrate off, are not the same product even though both read as "free" in a comparison grid.
Two cautions before you treat Flash as a free lunch. First, this is a vendor claim measured against z.ai's public offering, and we have flagged the whole cluster needs_native_reverify because our own price and latency numbers for non-DeepSeek families were captured through OpenRouter, not the native z.ai key. Second, "Flash" does not mean fast. When we measured GLM-4.7-Flash on a 200-token generation, it ran 26.24 seconds, the slowest single call in our entire cross-family test. Free, yes. Snappy, no.
According to Alibaba Cloud, the international Model Studio offering for Qwen includes a new-user free quota, commonly cited as roughly 1 million tokens valid for about 90 days on the Singapore endpoint. Treat the exact number as vendor-stated and reverify it before you plan around it; quotas like this change quietly and we have not re-pulled it on a native key.
The honest framing is that this is onboarding runway, not a permanent free tier. A million tokens is enough to evaluate Qwen3-Coder or Qwen-Plus thoroughly, ship a prototype, and decide whether to pay. It is not enough to run anything in production for free past the window. Plan the migration on day one.
According to DeepSeek API Docs, DeepSeek bills V4-Flash at $0.14 per million input tokens and $0.28 per million output, with a cached-input rate far below that. New accounts have historically received a sign-up credit, but the current amount is the one figure in this article we could not confirm to our own standard, so we are not going to quote a dollar number we cannot stand behind. If the credit matters to your decision, check it in the console at sign-up rather than trusting any list, including this one.
What is verified is that once the credit is gone, DeepSeek is among the cheapest paid options anywhere. For a fuller rate breakdown across families once the free runway ends, see the DeepSeek API pricing hub.
Documentation tells you a tier exists; it does not tell you what the model feels like to call. So we ran every family on the same small prompts and recorded latency and billed cost. The non-DeepSeek families here were measured via OpenRouter, not on the vendors' native free endpoints, so read these as capability and speed signals rather than as the exact bill a free-tier key would show. DeepSeek was measured on its official api.deepseek.com endpoint.
| Model | Latency (s) | Billed cost (USD) | Out tokens | Routing |
|---|---|---|---|---|
| GLM-4.7-Flash | 26.24 | $0.000104 | 200 | via OpenRouter |
| GLM-4.5-Air | 8.07 | $0.000174 | 200 | via OpenRouter |
| Qwen3-Coder | 1.50 | $0.0000656 | 58 | via OpenRouter |
| Qwen-Plus | 2.47 | $0.0000832 | 94 | via OpenRouter |
| DeepSeek V4-Flash | 1.96 | n/a (anchor) | 86 | official endpoint |
The numbers reframe the free-tier choice. GLM-4.7-Flash costs essentially nothing per call (just over one ten-thousandth of a dollar when we routed it), but its 26-second latency makes it a poor fit for anything interactive. What surprised us was the spread: the genuinely free model was also the slowest, while DeepSeek's paid V4-Flash answered in under two seconds. If your "free" requirement is really a "cheap and fast" requirement, the free tier may not be the right answer at all. Every figure above is restated from our cluster evidence pack so you can audit it.
Choose GLM-4.7-Flash or 4.5-Flash if you want a free tier that does not expire and you can tolerate slow responses, for example batch jobs, side projects, or learning. Choose Qwen's Singapore trial if you want a million tokens to evaluate a capable model and you accept that you will pay or leave after 90 days. Treat DeepSeek's credit as a small bonus on top of an already cheap paid plan, not as a free tier you can live on. Avoid assuming Kimi or MiniMax have a standing free tier; we found none to confirm.
Which Chinese LLM API is genuinely free? GLM is the only family here with a non-expiring free tier: GLM-4.7-Flash and GLM-4.5-Flash are offered free on z.ai. Qwen and DeepSeek offer trials or credits that run out, not standing free tiers.
How much does Qwen give new users for free? Alibaba Cloud's international Model Studio is commonly cited as granting roughly 1 million free tokens valid about 90 days on the Singapore endpoint. Treat this as a vendor-stated figure and reverify it at sign-up, since quotas change.
What is DeepSeek's free credit amount? DeepSeek has historically given new accounts a sign-up credit, but we could not verify the current amount to our own standard, so we do not quote a number. Check the console at sign-up. Paid V4-Flash is $0.14 per million input tokens.
Do Kimi and MiniMax have free API tiers? We found no standing free tier we could confirm for either Kimi (Moonshot) or MiniMax. Both are best treated as paid-only; check platform.moonshot.ai and platform.minimax.io for any current promotion before assuming free access.
Is the free GLM-4.7-Flash tier fast? No. In our cross-family test GLM-4.7-Flash took 26.24 seconds on a 200-token generation, the slowest call we measured. Free does not mean fast; for low latency a cheap paid tier like DeepSeek V4-Flash answered in under two seconds.
This is part of the best Chinese LLM API hub, which compares pricing, speed, and access across every family in this directory.
Author: Kevin Fan, Customer Success Manager at China LLM Directory, specializing in Chinese LLM ecosystem pricing and benchmarking. Last verified: 2026-06-26.
<!-- METADATA { "title": "Chinese LLM API Free Tiers, Compared (2026)", "slug": "chinese-llm-api-free-tiers", "meta_description": "Which Chinese LLM API is actually free? GLM-4.7-Flash is genuinely free on z.ai; Qwen gives ~1M trial tokens/90 days; DeepSeek's credit is unverified. 2026.", "focus_keyword": "chinese llm api free tier", "secondary_keywords": ["free chinese llm api", "glm free tier", "qwen free tokens", "deepseek free credit"], "tags": ["Free Tier", "API Pricing", "GLM", "Qwen", "DeepSeek"], "category": "Pricing", "cluster_id": "best-chinese-llm-api", "cluster_role": "micro", "hub_slug": "best-chinese-llm-api", "evidence_file": "clients/china-llm-aggregator/articles/best-chinese-llm-api-evidence.json", "needs_native_reverify": true, "verified_until": "2026-09-24", "author_name": "Kevin Fan", "author_title": "Customer Success Manager", "author_linkedin": "", "author_expertise": ["Chinese LLM ecosystem", "AI infrastructure pricing", "model benchmarking", "cross-border AI compliance"], "faq_pairs": [ {"q": "Which Chinese LLM API is genuinely free?", "a": "GLM is the only family here with a non-expiring free tier: GLM-4.7-Flash and GLM-4.5-Flash are offered free on z.ai. Qwen and DeepSeek offer trials or credits that run out, not standing free tiers."}, {"q": "How much does Qwen give new users for free?", "a": "Alibaba Cloud's international Model Studio is commonly cited as granting roughly 1 million free tokens valid about 90 days on the Singapore endpoint. Treat this as a vendor-stated figure and reverify it at sign-up, since quotas change."}, {"q": "What is DeepSeek's free credit amount?", "a": "DeepSeek has historically given new accounts a sign-up credit, but we could not verify the current amount to our own standard, so we do not quote a number. Check the console at sign-up. Paid V4-Flash is $0.14 per million input tokens."}, {"q": "Do Kimi and MiniMax have free API tiers?", "a": "We found no standing free tier we could confirm for either Kimi (Moonshot) or MiniMax. Both are best treated as paid-only; check platform.moonshot.ai and platform.minimax.io for any current promotion before assuming free access."}, {"q": "Is the free GLM-4.7-Flash tier fast?", "a": "No. In our cross-family test GLM-4.7-Flash took 26.24 seconds on a 200-token generation, the slowest call we measured. Free does not mean fast; for low latency a cheap paid tier like DeepSeek V4-Flash answered in under two seconds."} ], "external_links_used": [ {"url": "https://z.ai", "source_name": "z.ai", "claim": "GLM-4.7-Flash and GLM-4.5-Flash offered free on z.ai endpoint"}, {"url": "https://www.alibabacloud.com", "source_name": "Alibaba Cloud", "claim": "Qwen international Model Studio new-user free quota ~1M tokens/90 days on Singapore endpoint"}, {"url": "https://api-docs.deepseek.com/quick_start/pricing/", "source_name": "DeepSeek API Docs", "claim": "DeepSeek V4-Flash $0.14/M input, $0.28/M output; new-account credit exists, amount unverified"} ], "internal_links_used": [ {"url": "/blog/best-chinese-llm-api/", "anchor_text": "best Chinese LLM API hub", "type": "hub"}, {"url": "/blog/deepseek-api-pricing/", "anchor_text": "DeepSeek API pricing hub", "type": "cluster-hub"} ], "first_hand_evidence": { "source": "best-chinese-llm-api-evidence.json runs: glm-4.7-flash_general, glm-4.5-air_general, qwen3-coder_coding, qwen-plus_general, minimax deepseek_v4flash_coding_anchor", "measured": "GLM-4.7-Flash latency 26.244s billed $0.000103875 (200 out); GLM-4.5-Air 8.07s $0.00017374 (200 out); Qwen3-Coder 1.497s $0.0000656 (58 out); Qwen-Plus 2.47s $0.0000832 (94 out); DeepSeek V4-Flash 1.96s 86 out (official endpoint anchor); non-DeepSeek via OpenRouter", "captured": "2026-06-26" }, "images_status": "spec-only (not generated; FAL_API_KEY unset)", "images": [ {"position": "featured", "type": "generated", "prompt": "Clean editorial comparison chart of Chinese LLM free access tiers: GLM free forever, Qwen 90-day trial, DeepSeek sign-up credit, Kimi and MiniMax paid-only. Three-column traffic-light layout, blue and amber palette, minimal background, 16:9.", "alt": "Comparison chart of Chinese LLM API free tiers showing GLM as the only non-expiring free tier, Qwen as a 90-day trial, DeepSeek as a sign-up credit, and Kimi and MiniMax as paid-only"} ] } -->