At the cheap end DeepSeek wins on raw price: V4-Flash bills $0.14 in / $0.28 out on its official endpoint, while the closest cheap Qwen coding tier, qwen3-coder, runs $0.22 in / $1.80 out routed through OpenRouter; Qwen's flagship competes on capability, not cost. That single line hides an evidence asymmetry worth stating up front. The DeepSeek figures come from DeepSeek's own endpoint; the Qwen figures we cite are OpenRouter-routed, because Qwen is not reachable on a native key in our test environment. We keep the two clearly separated below so you budget against the number you will actually pay.
Start with the question most buyers actually ask: which is cheaper for the same job? The honest answer is "it depends which Qwen tier you mean." DeepSeek V4-Flash is the cheapest serious option in this comparison, and nothing in the Qwen lineup undercuts it on a blended per-token basis. But Qwen's flagship is not trying to win the price race; it is trying to win on capability, and its mid-tier coding model is close enough on price that latency can break the tie.
According to DeepSeek API Docs, DeepSeek V4-Flash bills $0.14 per million input tokens and $0.28 per million output tokens on the official api.deepseek.com endpoint, with an automatic prefix-cache discount on top. That is the anchor every Qwen tier gets measured against here.
| Model | Input ($/1M) | Output ($/1M) | Source |
|---|---|---|---|
| DeepSeek V4-Flash | $0.14 | $0.28 | official api.deepseek.com |
| qwen3-coder | $0.22 | $1.80 | OpenRouter-routed |
| qwen3-max | $0.78 | $3.90 | OpenRouter-routed |
| Qwen3-Max (official, 32K-128K input) | ~$2.40 | ~$12.00 | Alibaba Model Studio (pending re-verify) |
The output column is where the gap really opens. DeepSeek's $0.28 output rate is roughly 6.4x cheaper than qwen3-coder's $1.80 and about 14x cheaper than qwen3-max's $3.90. For any generation-heavy workload, code synthesis, long answers, summaries, output price dominates the bill, and that is DeepSeek's structural advantage.
Here is the part that surprises people. The OpenRouter-routed price for qwen3-max, $0.78 in / $3.90 out, is materially cheaper than Alibaba's own international flagship rate. According to Alibaba Cloud Model Studio pricing, the international Qwen3-Max tier is priced by input-size bracket, around $2.40 in / $12.00 out for the 32K-128K input window. We mark that official figure as pending native re-verification, but the direction is clear and it is a real buyer insight.
The practical upshot: an international buyer can pay roughly 3x less for the same Qwen3-Max model by routing through a third-party aggregator than by calling Alibaba's own International endpoint directly. That inverts the usual assumption that the vendor's first-party price is the floor. We do not hyperlink the routing provider here, but the routed number is the one our live test billed against, and we disclose its provenance every time we cite it.
Documentation gives you list prices. We wanted measured ones, so we sent a near-identical short coding prompt to each model and recorded real token counts, latency, and billed cost. The Qwen calls were routed via OpenRouter; the DeepSeek call hit its official endpoint. That asymmetry is the honest caveat, and it is exactly why we report the latency alongside the cost rather than pretending the two are perfectly matched.
When we called qwen3-coder via OpenRouter on a coding prompt, the run consumed 36 input and 58 output tokens, billed $0.0000656, and returned in 1.50 seconds. It was both the fastest and the cheapest of every Qwen tier we tested. The DeepSeek V4-Flash anchor on the same style of prompt used 27 input and 120 output tokens at 3.22 seconds; its cost was computed from the official $0.14/$0.28 rather than billed, because that call ran on DeepSeek's own endpoint.
| Run (coding prompt) | In tok | Out tok | Latency | Cost basis |
|---|---|---|---|---|
| qwen3-coder (via OpenRouter) | 36 | 58 | 1.50s | $0.0000656 billed |
| DeepSeek V4-Flash (official) | 27 | 120 | 3.22s | computed from $0.14/$0.28 |
Read that table carefully before drawing a winner. The two runs are not apples-to-apples: DeepSeek generated more than twice the output tokens, which inflates both its cost and its latency, and it ran through a different endpoint. What the run does show cleanly is that qwen3-coder returned in under half the wall-clock time on this prompt, 1.50s against 3.22s. Speed, not headline price, is where this Qwen tier earns its place against DeepSeek.
For a blended per-million-token bill on a generation-heavy workload, DeepSeek V4-Flash is the cheaper choice, and it is not close on the output side. Its $0.28 output rate against qwen3-coder's $1.80 means a workload producing far more output than input will favor DeepSeek by a wide margin, before you even count DeepSeek's automatic prefix cache.
According to DeepSeek API Docs, the V4-Flash cached input rate drops to $0.0028 per million, a 50x discount that fires automatically on repeated prefixes, which widens DeepSeek's cost lead further for chat workloads with stable system prompts. There is no comparable first-party cache discount in the Qwen figures we verified. For the full DeepSeek rate card and cache mechanics, see the DeepSeek API pricing cluster.
Choose DeepSeek V4-Flash when total cost per token is the deciding factor and your workload generates a lot of output. Choose Qwen when you need its flagship capability tier or want qwen3-coder's lower latency on short coding turns, and you have accepted the OpenRouter routing trade-off. The two are not really competing for the same buyer: one optimizes the bill, the other optimizes the ceiling.
Is Qwen or DeepSeek cheaper for API access? DeepSeek is cheaper at the low end. V4-Flash bills $0.14 in / $0.28 out officially, while the closest cheap Qwen tier, qwen3-coder, runs $0.22 in / $1.80 out via OpenRouter. DeepSeek's output rate is the decisive gap, roughly 6.4x cheaper than qwen3-coder's.
Why are your Qwen prices labeled "via OpenRouter"? Qwen is not reachable on a native Alibaba key in our test environment, so our first-hand Qwen measurements are OpenRouter-routed published and live-billed prices. Alibaba's official Model Studio rates are cited separately and marked pending native re-verification. We never present a routed price as Alibaba's official price.
Is the official Alibaba price the cheapest way to get Qwen3-Max? Not necessarily. The OpenRouter-routed rate of $0.78 in / $3.90 out is materially below Alibaba's international Qwen3-Max bracket of about $2.40 in / $12.00 out for the 32K-128K input window, so an international buyer can pay roughly 3x less by routing through an aggregator.
Was qwen3-coder faster than DeepSeek in your test? Yes, on our coding prompt. qwen3-coder returned in 1.50 seconds via OpenRouter against 3.22 seconds for the DeepSeek V4-Flash anchor. DeepSeek generated more than twice the output tokens on that run, so the latency comparison is not apples-to-apples, but qwen3-coder was the fastest tier we measured.
This is part of the Qwen API pricing hub, which covers every tier and routing path in one place.
Author: Kevin Fan, Customer Success Manager at China LLM Directory, specializing in Chinese LLM ecosystem pricing. Last verified: 2026-06-26.
<!-- METADATA { "title": "Qwen vs DeepSeek API Cost Compared (2026)", "slug": "qwen-vs-deepseek-cost", "meta_description": "DeepSeek V4-Flash ($0.14/$0.28) undercuts Qwen on raw price; qwen3-coder ($0.22/$1.80, OpenRouter) wins on latency. We billed both on a coding prompt. 2026-06.", "focus_keyword": "qwen vs deepseek api cost", "secondary_keywords": ["qwen vs deepseek pricing", "deepseek vs qwen api price", "qwen3-coder vs deepseek cost", "is qwen cheaper than deepseek"], "tags": ["Qwen", "DeepSeek", "API Pricing"], "category": "Pricing", "cluster_id": "qwen-api-pricing", "cluster_role": "micro", "hub_slug": "qwen-api-pricing", "evidence_file": "clients/china-llm-aggregator/articles/qwen-api-pricing-evidence.json", "needs_native_reverify": true, "verified_until": "2026-09-24", "author_name": "Kevin Fan", "author_title": "Customer Success Manager", "author_linkedin": "", "author_expertise": ["Chinese LLM ecosystem", "AI infrastructure pricing", "model benchmarking", "cross-border AI compliance"], "faq_pairs": [ {"q": "Is Qwen or DeepSeek cheaper for API access?", "a": "DeepSeek is cheaper at the low end. V4-Flash bills $0.14 in / $0.28 out officially, while the closest cheap Qwen tier, qwen3-coder, runs $0.22 in / $1.80 out via OpenRouter. DeepSeek's output rate is the decisive gap, roughly 6.4x cheaper than qwen3-coder's."}, {"q": "Why are your Qwen prices labeled via OpenRouter?", "a": "Qwen is not reachable on a native Alibaba key in our test environment, so our first-hand Qwen measurements are OpenRouter-routed published and live-billed prices. Alibaba's official Model Studio rates are cited separately and marked pending native re-verification. We never present a routed price as Alibaba's official price."}, {"q": "Is the official Alibaba price the cheapest way to get Qwen3-Max?", "a": "Not necessarily. The OpenRouter-routed rate of $0.78 in / $3.90 out is materially below Alibaba's international Qwen3-Max bracket of about $2.40 in / $12.00 out for the 32K-128K input window, so an international buyer can pay roughly 3x less by routing through an aggregator."}, {"q": "Was qwen3-coder faster than DeepSeek in your test?", "a": "Yes, on our coding prompt. qwen3-coder returned in 1.50 seconds via OpenRouter against 3.22 seconds for the DeepSeek V4-Flash anchor. DeepSeek generated more than twice the output tokens on that run, so the latency comparison is not apples-to-apples, but qwen3-coder was the fastest tier we measured."} ], "external_links_used": [ {"url": "https://api-docs.deepseek.com/quick_start/pricing/", "source_name": "DeepSeek API Docs - Pricing", "claim": "DeepSeek V4-Flash $0.14/$0.28 official; cache hit $0.0028/M 50x discount"}, {"url": "https://www.alibabacloud.com/help/en/model-studio/model-pricing", "source_name": "Alibaba Cloud Model Studio pricing", "claim": "Qwen3-Max international tier ~$2.40/$12.00 for 32K-128K input bracket (pending native re-verify)"} ], "internal_links_used": [ {"url": "/blog/qwen-api-pricing/", "anchor_text": "Qwen API pricing hub", "type": "hub"}, {"url": "/blog/deepseek-api-pricing/", "anchor_text": "DeepSeek API pricing", "type": "cross-cluster"} ], "first_hand_evidence": { "source": "qwen-api-pricing-evidence.json runs qwen3-coder_coding + deepseek_v4flash_coding_anchor", "measured": "qwen3-coder via OpenRouter: 36 in / 58 out, $0.0000656 billed, 1.50s; DeepSeek V4-Flash official: 27 in / 120 out, 3.22s, cost computed from $0.14/$0.28", "captured": "2026-06-26", "disclosure": "Qwen routed via OpenRouter; DeepSeek on official api.deepseek.com - asymmetry disclosed" }, "images_status": "spec-only (not generated; FAL_API_KEY unset)", "images": [ {"position": "featured", "type": "generated", "prompt": "Clean editorial side-by-side cost comparison of Qwen and DeepSeek API pricing, two columns labeled DeepSeek V4-Flash and Qwen, annotated with a coding-prompt latency result of 1.50s vs 3.22s. Blue and amber palette. Minimal background. 16:9.", "alt": "Side-by-side comparison of Qwen and DeepSeek API token pricing with a coding-prompt latency result of 1.50 seconds for qwen3-coder versus 3.22 seconds for DeepSeek V4-Flash"} ] } -->