Qwen API Pricing (2026)

Qwen3-Max routes through OpenRouter at $0.78 input and $3.90 output per 1M tokens, while Alibaba's own International endpoint runs about $2.40 / $12.00 on…

Fan Chuanyu's profile

Written by Fan Chuanyu

9 min read

Qwen3-Max routes through OpenRouter at $0.78 input and $3.90 output per 1M tokens, while Alibaba's own International endpoint runs about $2.40 / $12.00 on the same flagship, so an international buyer can pay roughly a third as much by not buying direct. That inversion is the most counterintuitive fact in Qwen pricing right now, and it sets the frame for every tier decision below. We measured the OpenRouter side first-hand on 2026-06-26; the official Alibaba figures are sourced and pending native re-verification.

This is the canonical landing page for Qwen API pricing aimed at buyers outside China. Qwen is Alibaba's family of open and hosted large language models, spanning a flagship (Qwen3-Max), a coding specialist (Qwen3-Coder), a balanced workhorse (Qwen-Plus), and several smaller MoE and dense variants. Below we map every tier to a price, show what each call actually cost when we ran it, and give you a decision matrix that no single sub-page repeats.

Qwen API pricing by tier (verified 2026-06)

Here is the rate card most buyers actually need before they pick an endpoint. Every Qwen row is the OpenRouter-routed published price, measured against the live OpenRouter catalog on 2026-06-26. The Alibaba official column is sourced from Model Studio and carries a re-verification caveat, so treat the two columns as different products, not a like-for-like discount.

Model (tier)OpenRouter-routed in/out ($/1M)Alibaba official in/out ($/1M)Best for
Qwen3-Max (flagship)$0.78 / $3.90~$2.40 / $12.00 (tiered)Hard reasoning, agents
Qwen3-Coder (480B-A35B)$0.22 / $1.80~$0.15 / $0.60Code generation
Qwen-Plus (balanced)$0.26 / $0.78~$0.40 / $1.20General chat at scale
Qwen3-32B (dense)$0.08 / $0.28variesCheap mid-size tasks
Qwen3.5-Flash$0.065 / $0.26variesHigh-volume short calls
Qwen3-235B-A22B-2507$0.09 / $0.10variesCheapest flagship-class MoE

The pattern worth noticing: for the flagship, the third-party route undercuts the vendor's own International endpoint, but for the coder and Qwen-Plus the official Alibaba rate is actually the cheaper of the two. The endpoint that wins flips depending on the tier, which is exactly why a single "Qwen is cheap" headline misleads.

Where the OpenRouter-vs-Alibaba gap comes from

According to Alibaba Cloud Model Studio pricing, the Qwen3-Max International rate is tiered by input size, landing near $2.40 input and $12.00 output per 1M tokens in the 32K to 128K bracket. That is roughly three times the $0.78 / $3.90 we see routed through OpenRouter, and the gap is structural rather than promotional. Third-party routers buy flagship capacity at a volume rate and pass a thinner margin, so the buyer who skips the direct relationship can come out ahead on the most expensive tier.

The reverse holds on the cheaper tiers. Qwen3-Coder's official Alibaba rate of about $0.15 / $0.60 sits below the $0.22 / $1.80 OpenRouter route, and Qwen-Plus official ($0.40 / $1.20) is mixed against its routed price ($0.26 / $0.78). The practical upshot is that there is no single "cheapest endpoint" for Qwen; the answer is per tier, and getting it wrong on output-heavy coding work is where buyers overspend.

We ran every tier through OpenRouter (first-hand evidence)

Published rate cards tell you the per-million price. They do not tell you what a real call costs or whether a "flash" model behaves the way its name promises. So we called each tier once through OpenRouter on 2026-06-26 and logged the billed cost, token counts, and latency. The full set lives in the evidence pack; here are the calls that change a buying decision.

The flagship was cheap and quick. Routed through OpenRouter, a Qwen3-Max general-knowledge call billed $0.00027534 for 38 input and 63 output tokens in 3.22 seconds. The coding specialist was cheaper still: Qwen3-Coder, served as qwen3-coder-480b-a35b, billed $0.0000656 for 36 input and 58 output tokens in 1.50 seconds, the fastest and cheapest call in the whole set.

Tier (via OpenRouter)In / out tokensBilled cost (USD)Latency
Qwen3-Max (general)38 / 63$0.000275343.22 s
Qwen3-Coder (coding)36 / 58$0.00006561.50 s
Qwen-Plus (general)38 / 94$0.00008322.47 s
Qwen3-32B (general)39 / 120$0.000036723.61 s
Qwen3.5-Flash (general)40 / 2831$0.0007386617.85 s

The Qwen3.5-Flash run is the one to sit with. On a two-sentence prompt of 40 input tokens it generated 2,831 output tokens and took 17.85 seconds, billing $0.00073866. That made the cheapest-per-token model the slowest and most expensive single call in our set. We re-read the response to be sure it was not an error; it was simply a verbose completion. The lesson is that "flash" names a price tier, not a guaranteed behavior, and an unbounded max_tokens on a chatty model can erase the per-token savings on a single request.

Cross-tier decision matrix (original synthesis)

No single sub-page in this cluster draws the whole picture, so here is the matrix we wish every Qwen buyer started with. Match your dominant workload to the tier and endpoint, not to the model with the lowest headline number.

Your workloadPick this tierPick this endpointWhy
Hard reasoning, agents, long contextQwen3-MaxOpenRouterRouted flagship ($0.78/$3.90) undercuts Alibaba International (~$2.40/$12.00)
Heavy code generationQwen3-CoderAlibaba officialOfficial ~$0.15/$0.60 beats routed $0.22/$1.80 on output
General chat at scaleQwen-PlusCompare bothRouted input ($0.26) is cheaper; official is mixed
High-volume, short, cost-criticalQwen3.5-Flash or 235B-A22BOpenRouterLowest per-token, but cap max_tokens to avoid runaway calls
US-based, needs free quotaAny tierAlibaba Singapore~1M free tokens for 90 days on Model Studio International

The rule the matrix encodes: cap output on the cheap tiers, route the flagship rather than buying it direct, and buy the coder direct rather than routing it. Choose OpenRouter when the flagship dominates your spend. Choose Alibaba official when coding output volume dominates, or when you want the Singapore free quota to prototype.

Is Qwen available to US and international buyers?

Yes, through two paths, and the access question is separate from the price question. According to Alibaba Cloud Model Studio pricing, Qwen is offered on Model Studio International with a Singapore endpoint that includes roughly 1M free tokens for 90 days, plus a US (Virginia) Global deployment that carries no free quota. US developers can therefore reach Qwen directly from Alibaba, in addition to the third-party routes.

That matters for buyers who assumed a Chinese-model meant no Western data residency option. A Virginia deployment puts Qwen inference on US soil, which is the detail compliance teams usually ask about first. The free Singapore quota is the cheapest way to benchmark the tiers yourself before committing to either endpoint.

How Qwen compares to DeepSeek and the US flagships

Qwen does not live in a vacuum, and the cross-vendor frame is where the directory earns its keep. According to DeepSeek API Docs, DeepSeek V4-Flash bills $0.14 input and $0.28 output per 1M tokens on its official endpoint, which undercuts every Qwen tier on output except the 235B-A22B MoE. For our DeepSeek anchor we called the official api.deepseek.com endpoint, while Qwen ran through OpenRouter, so read that comparison as endpoint-asymmetric rather than apples-to-apples.

Against the US flagships the gap favors Qwen. According to DeepSeek API Docs, DeepSeek output is $0.28 per 1M; GPT-4o sits at $2.50 / $10.00 and Claude Sonnet 4.6 at $3 / $15. Routed Qwen3-Max at $3.90 output undercuts both Western flagships while staying in the same reasoning class, the value case that keeps Chinese models on the shortlist.

Explore the full Qwen pricing cluster

Each sub-page below answers one buyer question in depth. Use this table as the cluster map.

Sub-topicPage
Qwen3-Max per-token pricingQwen3-Max pricing
Qwen3-Coder pricingQwen3-Coder pricing
Qwen-Plus pricingQwen-Plus pricing
Qwen vs DeepSeek costQwen vs DeepSeek cost
Qwen vs GPT-4o costQwen vs GPT-4o cost
Cheapest Qwen modelCheapest Qwen model
OpenRouter vs official Qwen pricingOpenRouter vs official pricing
US availabilityIs Qwen available in the US
Cost calculatorQwen API cost calculator
Coder vs DeepSeek coding costQwen3-Coder vs DeepSeek coding cost

FAQ

How much does the Qwen API cost? Routed through OpenRouter, Qwen3-Max bills $0.78 input and $3.90 output per 1M tokens, Qwen3-Coder $0.22 / $1.80, and Qwen-Plus $0.26 / $0.78. Alibaba's own International endpoint prices the flagship higher, near $2.40 / $12.00, so the cheapest endpoint depends on the tier.

Is Qwen cheaper through OpenRouter or directly from Alibaba? It depends on the tier. On the Qwen3-Max flagship the OpenRouter route ($0.78/$3.90) undercuts Alibaba International ($2.40/$12.00). On Qwen3-Coder the official Alibaba rate ($0.15/$0.60) is cheaper than the routed $0.22/$1.80. There is no single winner.

Can US developers use the Qwen API? Yes. Alibaba Cloud Model Studio International offers a Singapore endpoint with roughly 1M free tokens for 90 days and a US (Virginia) Global deployment with no free quota. Qwen is also reachable through third-party routers from outside China.

What was the most surprising thing in your testing? Qwen3.5-Flash, the cheapest per-token tier, produced 2,831 output tokens on a two-sentence prompt and took 17.85 seconds, making it the slowest and most expensive single call we measured via OpenRouter. A low per-token price does not cap a single call's cost; your max_tokens does.

How does Qwen compare to DeepSeek on price? DeepSeek V4-Flash bills $0.14 / $0.28 on its official endpoint, cheaper on output than every Qwen tier except the 235B-A22B MoE. We called DeepSeek on its official endpoint and Qwen via OpenRouter, so treat that comparison as endpoint-asymmetric.

Is Qwen cheaper than GPT-4o? Yes, by a wide margin. GPT-4o costs $2.50 / $10.00 per 1M tokens. Routed Qwen3-Max at $0.78 / $3.90 undercuts it on both input and output while staying in a comparable reasoning class.

Which Qwen tier is cheapest overall? By published per-token rate, Qwen3-235B-A22B-2507 at $0.09 / $0.10 via OpenRouter is the cheapest flagship-class option, and Qwen3.5-Flash at $0.065 input is cheapest on input. Cap output to keep a single call from running away.


Author: Kevin Fan, Customer Success Manager at China LLM Directory, specializing in Chinese LLM ecosystem pricing and cross-border AI access. Last verified: 2026-06-26.

<!-- METADATA { "title": "Qwen API Pricing: Complete Buyer Guide (2026)", "slug": "qwen-api-pricing", "meta_description": "Qwen3-Max routes via OpenRouter at $0.78/$3.90 per 1M, below Alibaba's own ~$2.40/$12 International rate. Tier prices, first-hand tests, decision matrix.", "focus_keyword": "qwen api pricing", "secondary_keywords": ["qwen3-max pricing", "qwen api cost", "alibaba qwen pricing", "qwen openrouter price", "qwen vs deepseek cost"], "tags": ["Qwen", "Alibaba Cloud", "API Pricing", "OpenRouter"], "category": "Pricing", "cluster_id": "qwen-api-pricing", "cluster_role": "hub", "evidence_file": "clients/china-llm-aggregator/articles/qwen-api-pricing-evidence.json", "needs_native_reverify": true, "verified_until": "2026-09-24", "author_name": "Kevin Fan", "author_title": "Customer Success Manager", "author_linkedin": "", "author_expertise": ["Chinese LLM ecosystem", "AI infrastructure pricing", "model benchmarking", "cross-border AI compliance"], "faq_pairs": [ {"q": "How much does the Qwen API cost?", "a": "Routed through OpenRouter, Qwen3-Max bills $0.78 input and $3.90 output per 1M tokens, Qwen3-Coder $0.22 / $1.80, and Qwen-Plus $0.26 / $0.78. Alibaba's own International endpoint prices the flagship higher, near $2.40 / $12.00, so the cheapest endpoint depends on the tier."}, {"q": "Is Qwen cheaper through OpenRouter or directly from Alibaba?", "a": "It depends on the tier. On the Qwen3-Max flagship the OpenRouter route ($0.78/$3.90) undercuts Alibaba International (~$2.40/$12.00). On Qwen3-Coder the official Alibaba rate (~$0.15/$0.60) is cheaper than the routed $0.22/$1.80. There is no single winner."}, {"q": "Can US developers use the Qwen API?", "a": "Yes. Alibaba Cloud Model Studio International offers a Singapore endpoint with roughly 1M free tokens for 90 days and a US (Virginia) Global deployment with no free quota. Qwen is also reachable through third-party routers from outside China."}, {"q": "What was the most surprising thing in your testing?", "a": "Qwen3.5-Flash, the cheapest per-token tier, produced 2,831 output tokens on a two-sentence prompt and took 17.85 seconds, making it the slowest and most expensive single call we measured via OpenRouter. A low per-token price does not cap a single call's cost; your max_tokens does."}, {"q": "How does Qwen compare to DeepSeek on price?", "a": "DeepSeek V4-Flash bills $0.14 / $0.28 on its official endpoint, cheaper on output than every Qwen tier except the 235B-A22B MoE. We called DeepSeek on its official endpoint and Qwen via OpenRouter, so treat that comparison as endpoint-asymmetric."}, {"q": "Is Qwen cheaper than GPT-4o?", "a": "Yes, by a wide margin. GPT-4o costs $2.50 / $10.00 per 1M tokens. Routed Qwen3-Max at $0.78 / $3.90 undercuts it on both input and output while staying in a comparable reasoning class."}, {"q": "Which Qwen tier is cheapest overall?", "a": "By published per-token rate, Qwen3-235B-A22B-2507 at $0.09 / $0.10 via OpenRouter is the cheapest flagship-class option, and Qwen3.5-Flash at $0.065 input is cheapest on input. Cap output to keep a single call from running away."} ], "external_links_used": [ {"url": "https://www.alibabacloud.com/help/en/model-studio/model-pricing", "source_name": "Alibaba Cloud Model Studio pricing", "claim": "Qwen3-Max International tiered pricing ~$2.40/$12.00; Singapore free quota ~1M tokens 90 days; US Virginia Global deployment"}, {"url": "https://api-docs.deepseek.com/quick_start/pricing/", "source_name": "DeepSeek API Docs – Pricing", "claim": "DeepSeek V4-Flash $0.14/$0.28 official endpoint, comparison anchor for Qwen"} ], "internal_links_used": [ {"url": "/blog/qwen3-max-pricing/", "anchor_text": "Qwen3-Max pricing", "type": "micro"}, {"url": "/blog/qwen3-coder-pricing/", "anchor_text": "Qwen3-Coder pricing", "type": "micro"}, {"url": "/blog/qwen-plus-pricing/", "anchor_text": "Qwen-Plus pricing", "type": "micro"}, {"url": "/blog/qwen-vs-deepseek-cost/", "anchor_text": "Qwen vs DeepSeek cost", "type": "micro"}, {"url": "/blog/qwen-vs-gpt-4o-cost/", "anchor_text": "Qwen vs GPT-4o cost", "type": "micro"}, {"url": "/blog/cheapest-qwen-model/", "anchor_text": "Cheapest Qwen model", "type": "micro"}, {"url": "/blog/qwen-openrouter-vs-official-pricing/", "anchor_text": "OpenRouter vs official pricing", "type": "micro"}, {"url": "/blog/is-qwen-available-in-us/", "anchor_text": "Is Qwen available in the US", "type": "micro"}, {"url": "/blog/qwen-api-cost-calculator/", "anchor_text": "Qwen API cost calculator", "type": "micro"}, {"url": "/blog/qwen3-coder-vs-deepseek-coding-cost/", "anchor_text": "Qwen3-Coder vs DeepSeek coding cost", "type": "micro"} ], "first_hand_evidence": { "source": "qwen-api-pricing-evidence.json runs (via OpenRouter; DeepSeek anchor official)", "measured": "qwen3-max general 38in/63out $0.00027534 3.22s; qwen3-coder (qwen3-coder-480b-a35b) 36in/58out $0.0000656 1.50s; qwen-plus 38in/94out $0.0000832 2.47s; qwen3-32b 39in/120out $0.00003672 3.61s; qwen3.5-flash 40in/2831out $0.00073866 17.85s", "captured": "2026-06-26", "disclosure": "Qwen measured via OpenRouter; official Alibaba prices pending native re-verification (needs_native_reverify)" }, "images_status": "spec-only (not generated; FAL_API_KEY unset)", "images": [ {"position": "featured", "type": "generated", "prompt": "Clean editorial comparison diagram of Qwen API tiers, showing Qwen3-Max routed via OpenRouter at $0.78/$3.90 sitting below Alibaba's own International endpoint at ~$2.40/$12.00, with smaller tiers (Coder, Plus, Flash) stacked below. Blue and amber palette, minimal background, 16:9.", "alt": "Diagram comparing Qwen API tier prices, showing Qwen3-Max routed via OpenRouter at $0.78 input and $3.90 output undercutting Alibaba's own International endpoint at about $2.40 and $12.00 per million tokens"} ] } -->

Share: