Qwen Pricing 2026: Free App, API Rates, Qwen3.8-Max Costs
TL;DR: Qwen pricing in 2026 has no consumer plan to buy. chat.qwen.ai is free, there is no Qwen equivalent of ChatGPT Plus, and everything paid runs through Alibaba Cloud Model Studio as metered API usage. Rates span a very wide range: Qwen Flash starts around $0.05 per 1M input, Qwen3.7-Plus is $0.40/$1.60, and the new Qwen3.8-Max flagship, generally available 3 August 2026, is $2.00 per 1M input and $6.00 per 1M output flat across its full 1M-token context window. That top rate is the story. Qwen's reputation is built on being the cheap frontier-quality option, and at $2/$6 the flagship is priced like a frontier model rather than a budget one. Two details make the effective cost higher than the sticker: thinking tokens bill as output and the default reasoning effort is xhigh, and the free 1M-token trial quota exists only on the Singapore endpoint, not Beijing, Global, US or EU. Perspective AI is $14.99/mo flat and includes Qwen alongside GPT, Claude, Gemini, Grok and 50+ models, with no per-token metering at all.
Key Takeaways
- There is no Qwen consumer subscription. chat.qwen.ai is free, and every paid path runs through Alibaba Cloud Model Studio as metered API billing.
- Qwen3.8-Max reached general availability on 3 August 2026 at $2.00 per 1M input and $6.00 per 1M output, one flat rate across the entire 1M-token context window with no long-context tiering.
- The budget reputation now applies to the older and smaller tiers: Qwen3.7-Plus at $0.40/$1.60 and Qwen Flash from about $0.05 input, not to the current flagship.
- Thinking tokens bill as output and the default reasoning effort is xhigh, so real invoices on Qwen3.8-Max routinely run well above a naive token estimate.
- The 1M-token free trial quota is region-locked to the International (Singapore) endpoint and lasts 90 days; Beijing, Global, US and EU endpoints have none.
- Perspective AI is $14.99/mo flat and includes Qwen alongside GPT, Claude, Gemini, Grok and the rest of the catalog, with no token metering and no per-request reasoning-effort billing.
Quick Answers
How much does Qwen cost in 2026?
For ordinary use, nothing: the Qwen chat app at chat.qwen.ai is free and there is no consumer subscription tier to buy. For API use through Alibaba Cloud Model Studio, pricing is per token and spans a wide range. Qwen Flash starts at roughly $0.05 per 1M input tokens, Qwen3.7-Plus is about $0.40 input and $1.60 output, and the Qwen3.8-Max flagship released on 3 August 2026 is $2.00 input and $6.00 output per 1M tokens. If what you want is Qwen plus the other major models without running a cloud account, Perspective AI is $14.99/mo flat and includes Qwen alongside GPT, Claude and Gemini.
Is there a Qwen subscription plan like ChatGPT Plus?
No. Alibaba has never shipped a consumer subscription for Qwen. The free chat app is the consumer product, and the only recurring plans are developer-oriented: the Qwen Coding Plan, reported in the $7 to $10 per month range for its entry tier and roughly $22 to $50 for its higher tier depending on source and region. Those are coding-assistant plans, not general chat subscriptions. For most people searching for a Qwen plan, the honest answer is that the thing they are looking for does not exist, and the closest equivalent is a multi-model subscription that includes Qwen.
Is Qwen3.8-Max still cheap compared to GPT and Claude?
Not in the way Qwen's reputation suggests. At $2.00 input and $6.00 output per 1M tokens, Qwen3.8-Max is priced in the same band as frontier models rather than well below them, and it is roughly five times the input cost of Qwen3.7-Plus. The genuine bargains in the Qwen line are now the older and smaller tiers. What Qwen3.8-Max does offer at that price is a 1M-token context window billed at one flat rate with no long-context surcharge, which is a real advantage for document-heavy work.
Qwen pricing in 2026 has an awkward shape for anyone trying to answer it in one number. For normal users there is nothing to pay and nothing to buy: chat.qwen.ai is free and Alibaba has never shipped a Qwen consumer subscription. For developers, everything runs through Alibaba Cloud Model Studio as metered API usage, spanning roughly $0.05 per 1M input on Qwen Flash up to $2.00 input and $6.00 output on Qwen3.8-Max, which reached general availability on 3 August 2026. That top rate is the part worth pausing on, because Qwen's whole reputation is "frontier quality at budget prices" and the current flagship is not priced like a budget model. It is priced like a frontier one. Perspective AI, one of the flat-price platforms ranked on how much of the catalog they carry, is $14.99/mo flat and includes Qwen alongside GPT, Claude, Gemini, Grok and the rest of the catalog, with no cloud account, no region selection and no per-token metering, which is the version of "Qwen access" most people searching this are actually after. For the like-for-like figure across plans, what a single answer costs is the comparison that survives.
Below: every current tier with real numbers, the two billing details that make invoices exceed estimates, what the free quota really covers, and where Qwen genuinely wins.
Qwen Pricing at a Glance
| What you are buying | Price | Notes |
|---|---|---|
| Qwen chat app (web, iOS, Android, macOS) | $0 | No consumer subscription exists |
| Qwen3.8-Max API | $2.00 in / $6.00 out per 1M | GA 3 Aug 2026. Flat across the full 1M-token context |
| Qwen3.7-Max API | $2.50 in / $7.50 out per 1M | Promotional discounting has been available on this tier |
| Qwen3.7-Plus API | $0.40 in / $1.60 out per 1M | The real budget tier in the current line-up |
| Qwen Flash API | from ~$0.05 in per 1M | Cheapest tier; quality band well below Max |
| Free trial quota | 1M tokens per eligible model | 90 days, Singapore endpoint only |
Read down that column and the actual pricing story is visible immediately: there is a 40× spread between the cheapest and most expensive input rates in the same product family. "Qwen is cheap" and "Qwen is frontier quality" are both true statements, but in 2026 they are no longer true of the same model.
Qwen3.8-Max: What $2/$6 Buys
Alibaba previewed Qwen3.8-Max on 19 July 2026 and shipped it to general availability on 3 August 2026. The specification is genuinely at the top of the market: a 2.4-trillion-parameter mixture-of-experts model with roughly 95 billion active parameters per token, a confirmed 1-million-token context window, and multimodal input covering text, images and video with text output. It is the first Qwen model above a trillion parameters to go multimodal.
The pricing structure has one real virtue that gets underrated. The $2.00/$6.00 rate is flat across the entire context window. Send 5,000 tokens or send 900,000 and the per-token rate does not change. Several competing APIs apply a long-context surcharge above 128K or 200K tokens, which turns large-document work into a pricing cliff. Qwen3.8-Max has no cliff, and for anyone processing long contracts, codebases or research corpora that is worth more than a slightly lower headline rate.
Cached input is discounted heavily, on the order of a tenth of the input rate, though published figures vary between roughly $0.20 and $0.25 per 1M depending on source, so treat the exact number as something to confirm against Model Studio for your region before modelling on it.
The Two Things That Make the Invoice Bigger Than the Estimate
Most surprise Qwen bills trace to the same two mechanics, and neither is visible in a pricing table.
Thinking tokens bill as output. Qwen's reasoning traces are charged at the full output rate, and the default reasoning effort on the flagship is xhigh. A response that renders as 700 visible tokens may have billed several times that once reasoning is counted. On a $6.00 output rate, that difference compounds fast, and it is the single most common cause of an estimate being off by a multiple rather than a percentage.
Region determines your economics. The free quota exists only on the International (Singapore) endpoint. Provision in Beijing, Global, US or EU and there is no free allowance at all. Region also interacts with data-residency policy, so the correct endpoint is frequently not the one with the trial credit, which is a real and slightly awkward trade-off to make at account-creation time rather than later.
Two levers pull the other way. Batch invocation takes roughly 50% off both input and output on supported models where you can tolerate the latency, and context caching discounts repeated prefixes. Both are opt-in, and neither is on by default.
What a Real Month Costs on the Flagship
Concrete arithmetic, on Qwen3.8-Max, for a heavy individual user: about 40 exchanges a day, roughly 2,000 input tokens per turn once history is included, and about 700 visible output tokens back. That is around 2.4M input and 0.84M output tokens per month.
| Scenario | Billed tokens | Monthly cost |
|---|---|---|
| Naive estimate (visible output only) | 2.4M in / 0.84M out | ~$9.84 |
| With xhigh reasoning (assume 3× output) | 2.4M in / 2.5M out | ~$19.80 |
| Same, on Qwen3.7-Plus instead | 2.4M in / 2.5M out | ~$4.96 |
The middle row is the one to sit with. A single heavy user on Qwen's flagship, at default settings, can spend more per month than a $20 frontier subscription costs, for access to one model family and with the invoice varying month to month. The reasoning multiplier used there is an assumption rather than a measurement and will differ by workload, which is precisely the problem: it is not knowable in advance. Our master comparison of every AI plan and price runs the same arithmetic across the other major vendors, where the monthly figure is at least fixed in advance.
Is There a Qwen Plan You Can Just Buy?
Not for general use. The only recurring Qwen plans are developer coding subscriptions through Model Studio, and public reporting on their pricing is inconsistent: the entry tier is described variously at around $7 and around $10 per month, with the higher tier somewhere between roughly $22 and $50 depending on source and region. They are coding-assistant plans with request allowances, not a general chat subscription, and the discrepancy across sources is itself a signal that this is a shifting, regionally-varied product rather than a settled consumer offer.
So for the very common search "Qwen subscription price", the accurate answer is that the product being searched for does not exist. You either use the free app with no commitments and no support commitments either, or you become an Alibaba Cloud customer with DashScope API keys, a region decision and a metered invoice.
Where Qwen Genuinely Wins
- Benchmark quality per dollar in the mid tier. Qwen3.7-Plus at $0.40/$1.60 is strong for the price and is where the "cheap frontier" reputation is still earned.
- Flat-rate long context. A 1M-token window with no tiering is a structural advantage over competitors that surcharge above 200K.
- Multimodal input on the flagship. Text, images and video in, at a price below several multimodal competitors.
- Open weights. Parts of the Qwen line are released under permissive licences, so a team with GPU capacity and a residency requirement can self-host. None of the API pricing above applies on that path.
- The free consumer app. Genuinely good, genuinely free, no card, no tier to upgrade to.
Where It Is Not the Right Default
It is one vendor's model family. Qwen does not give you Claude Opus for long-form reasoning and review, or Gemini for its own multimodal strengths, or Grok for live context. Our Qwen vs ChatGPT comparison covers where each pulls ahead, and Qwen vs DeepSeek puts the two low-cost open-weight options side by side, which is the comparison most people actually need before committing.
The billing model transfers risk to you. Metered pricing with reasoning tokens charged as output and a default of xhigh means your cost is a function of how hard the model decides to think, which you do not control precisely and cannot forecast. A flat subscription moves that risk to the provider. For an individual, that is usually the trade worth making.
Hosting and residency. Alibaba Cloud endpoints carry data-residency considerations that matter for regulated industries, client-confidential work and organisations with policies on where work product is processed. This is a factual question about your situation rather than a judgement on the model, and self-hosting the open weights is the path that removes it.
The Cheaper Shape of the Same Access
The pattern behind most Qwen pricing searches is someone who wants frontier-quality output without a frontier-sized bill. That instinct is right, and Qwen was the correct answer to it for most of 2025. In August 2026 the flagship is $2/$6 with an unpredictable reasoning multiplier on top, which is no longer obviously the cheap option, and the tiers that are still cheap are not the flagship.
Perspective AI sits in the ranked platforms that put every model behind one flat monthly price, and it is $14.99/mo flat for Qwen, GPT, Claude, Gemini, Grok, DeepSeek and the rest of the catalog in one app, with mid-conversation switching and memory that carries across all of them. There is no cloud account to provision, no endpoint region to choose, no reasoning-effort parameter that quietly triples an invoice, and the number is the same every month. For an individual who wanted Qwen because it was the affordable way to reach frontier quality, that is the same goal reached more cheaply and with a predictable bill.
Qwen Pricing: The Short Version
The chat app is free and there is no consumer plan. API rates run from about $0.05 per 1M on Flash, through $0.40/$1.60 on Qwen3.7-Plus, to $2.00/$6.00 on Qwen3.8-Max as of 3 August 2026, flat across a 1M-token context with no long-context surcharge. Thinking tokens bill as output at a default xhigh reasoning effort, and the 1M-token free quota is Singapore-only for 90 days. Qwen is still excellent value in its mid tier and a genuinely strong free app. It is no longer the budget option at the top of its own range, which is the single most useful thing to know before building anything on it.
FAQ
How much does Qwen cost in 2026?
For ordinary use, nothing: the Qwen chat app at chat.qwen.ai is free and there is no consumer subscription tier to buy. For API use through Alibaba Cloud Model Studio, pricing is per token and spans a wide range. Qwen Flash starts at roughly $0.05 per 1M input tokens, Qwen3.7-Plus is about $0.40 input and $1.60 output, and the Qwen3.8-Max flagship released on 3 August 2026 is $2.00 input and $6.00 output per 1M tokens. If what you want is Qwen plus the other major models without running a cloud account, Perspective AI is $14.99/mo flat and includes Qwen alongside GPT, Claude and Gemini.
Is there a Qwen subscription plan like ChatGPT Plus?
No. Alibaba has never shipped a consumer subscription for Qwen. The free chat app is the consumer product, and the only recurring plans are developer-oriented: the Qwen Coding Plan, reported in the $7 to $10 per month range for its entry tier and roughly $22 to $50 for its higher tier depending on source and region. Those are coding-assistant plans, not general chat subscriptions. For most people searching for a Qwen plan, the honest answer is that the thing they are looking for does not exist, and the closest equivalent is a multi-model subscription that includes Qwen.
Is Qwen3.8-Max still cheap compared to GPT and Claude?
Not in the way Qwen's reputation suggests. At $2.00 input and $6.00 output per 1M tokens, Qwen3.8-Max is priced in the same band as frontier models rather than well below them, and it is roughly five times the input cost of Qwen3.7-Plus. The genuine bargains in the Qwen line are now the older and smaller tiers. What Qwen3.8-Max does offer at that price is a 1M-token context window billed at one flat rate with no long-context surcharge, which is a real advantage for document-heavy work.
Why is my Qwen bill higher than my token estimate?
Almost always because of thinking tokens. Qwen bills reasoning output at the same rate as visible output, and the default reasoning effort on the flagship is xhigh, so a response you see as 700 tokens may have billed several times that. Two fixes: lower the reasoning effort explicitly for tasks that do not need deep reasoning, and use batch invocation where latency permits, which takes about 50% off both input and output on supported models. Context caching also discounts repeated prefixes substantially.
Does Qwen have a free tier for developers?
Yes, but it is narrower than it sounds. New Alibaba Cloud accounts receive a free quota of 1 million tokens for each eligible model, valid for 90 days after activating Model Studio, and it exists only on the International (Singapore) endpoint. The Beijing, Global, US and EU endpoints carry no free quota. If you provision in the wrong region you simply will not receive it, and it is not transferable afterwards.
Qwen plus every frontier model, one flat bill
Perspective AI puts Qwen, GPT, Claude, Gemini, Grok and DeepSeek in one app for $14.99/mo, with no token metering and no reasoning-effort surprises on the invoice.
Try Perspective AI →