Pricing

Qwen Pricing 2026

Alibaba · qwen3.8-max

Qwen has no consumer subscription to buy. The chat app is free at $0, and everything paid is metered API usage through Alibaba Cloud Model Studio, from $0.05 per 1M input tokens on the cheapest tier to $2.00 input and $6.00 output on the flagship.

  • Chat app: chat.qwen.ai, with no consumer subscription of any kind, $0
  • qwen3.8-max: the flagship, one flat rate across the full one million token window, $2.00 in / $6.00 out per 1M
  • qwen3.7-plus: the real budget tier, but banded: it steps to $1.20 and $4.80 above 256,000 tokens, $0.40 in / $1.60 out per 1M
  • qwen-flash: the cheapest line on the card, up to 256,000 tokens, $0.05 in / $0.40 out per 1M
  • qwen3.5-flash: flash pricing held flat across a full million tokens, $0.10 in / $0.40 out per 1M
  • Cached input: expressed as a formula: a cache hit costs about 10% of the standard input rate, about 10% of input
  • Free trial quota: reported as a million tokens per model for 90 days, on the Singapore endpoint only, $0, region-locked

Alibaba Cloud refused a direct automated request on 2026-08-19, so every rate here was read through a read-through fetch of that same first-party URL rather than from a comparison site, and the flagship figures are independently corroborated. The card carried at least two active promotional discounts on the day it was read, so re-check before budgeting: Alibaba Cloud Model Studio pricing.

Free tier
$0
  • Free at $0 at chat.qwen.ai with no consumer subscription of any kind. Paid access is metered API usage through Alibaba Cloud Model Studio
  • Free-tier score: 4/5
What's included
Free
  • Full access to qwen3.8-max
  • Context window: 1M tokens
  • Voice mode
  • Image generation
  • Web search
  • File upload
API pricing
$2.00per 1M input tokens
$6.00per 1M output tokens
~$3.20est. per 1M in + 200K out
Context window
1M
Free tier
4/5
Open weights
Yes

There is no Qwen plan to buy

The most useful thing to say about Qwen pricing is that the product most people are searching for does not exist. Alibaba has never shipped a consumer subscription for Qwen. There is no Plus tier, no Pro tier and no monthly plan: the chat app is free, and everything paid runs through Alibaba Cloud Model Studio as metered API usage with a cloud account, an endpoint region and a per-token invoice attached.

So the honest answer splits in two. If you want to use Qwen, it costs nothing. If you want to build on Qwen, it costs whatever your tokens cost, and the range is far wider than this vendor's reputation suggests.

The Model Studio rate card

Read from Alibaba Cloud's own Model Studio pricing documentation on 2026-08-19. Rates are per million tokens in US dollars.

ModelInputOutputContext bandNotes
qwen3.8-max $2.00 $6.00 0 to 1M, one rate The flagship. Genuinely flat across the full window
qwen3.7-max $2.50 list $7.50 list n/a Carrying a 50% discount on the card when read, so $1.25 and $3.75 in effect
qwen3.7-plus $0.40 $1.60 0 to 256K The real budget tier. A limited-time discount was noted on the card
qwen3.7-plus, long context $1.20 $4.80 256K to 1M A 3x step at the boundary. See the section below
qwen-flash $0.05 $0.40 0 to 256K The cheapest line on the card
qwen3.5-flash $0.10 $0.40 0 to 1M Flat across the full window at a flash-tier price
qwen3.6-flash $0.25 $1.50 n/a Newer flash variant, priced above the older ones

Read down the input column and the pricing story is visible immediately: there is a 40-fold spread between the cheapest and the most expensive input rate inside a single product family. Both halves of this vendor's reputation are true, and in 2026 they are no longer true of the same model. The cheap Qwen and the frontier-class Qwen are different line items.

One provenance note, because it changes how much weight these numbers should carry. Alibaba's documentation refused a direct automated request on the verification date and was read through a read-through fetch of the same first-party URL. That is the vendor's own page rather than a comparison table, and the flagship figures match an independently researched draft that reached them separately. Treat the card as current and dated, and re-read it before committing a budget, because it carried at least two active promotional discounts on the day it was read and promotional rates expire.

The long-context cliff most Qwen write-ups miss

Qwen is widely recommended for long-document work on the grounds that it charges one flat rate across a very large context window, with no surcharge of the kind several competitors apply above 128K or 200K tokens. That is true of the flagship and it is not true of the tier most people would actually choose.

qwen3.7-plus is banded. Up to 256,000 tokens it is $0.40 and $1.60. Above 256,000 tokens and up to a million, the same model on the same card is $1.20 and $4.80. That is a three-fold step, and it lands precisely where a long-document workload lives. A pipeline that averages 200K tokens a request and occasionally spikes over the boundary does not get proportionally more expensive at the margin, it changes rate.

Two tiers do hold one rate across the whole window: qwen3.8-max at $2.00 and $6.00, and qwen3.5-flash at $0.10 and $0.40. If flat-rate long context is the reason you are here, those are the two lines that deliver it, and the second one is a twentieth of the price of the first. Check which band your median request falls in before choosing on the headline rate.

Why the invoice exceeds the estimate

Two mechanics account for most surprise Qwen bills, and neither is visible in a rate table.

Reasoning output bills as output. Tokens the model spends thinking before it answers are charged at the full output rate, so a response that renders as 700 tokens can bill several times that. On a $6.00 output rate the gap compounds quickly, and it is the reason an estimate is usually wrong by a multiple rather than by a percentage. The reasoning effort is a request parameter, so this is controllable, but only if you set it deliberately: check what your client library sends by default before assuming it is low, and benchmark at the effort level you will actually ship rather than at the minimum.

Region decides your economics before you write any code. The free trial quota, widely reported as a million tokens per eligible model for 90 days, exists only on the International endpoint served from Singapore. Provision in another region and there is no free allowance at all, and it is not transferable afterwards. Region also interacts with data-residency policy, so the correct endpoint for a given organisation is frequently not the one carrying the trial credit. That trade-off has to be made at account creation, which is the worst possible time to discover it.

Two levers pull the other way, and both are opt-in. Batch invocation takes roughly half off input and output on supported models where the latency is tolerable. Context caching is expressed on the card as a formula rather than a flat price: writing to the cache costs about 125% of the standard input rate, and a cache hit costs about 10% of it. That is a tenfold reduction on the repeated part of a prompt, which for a prefix-heavy workload is a larger saving than any model downgrade.

What a month costs on the flagship

Concrete arithmetic on qwen3.8-max, for a heavy individual: roughly 40 exchanges a day, about 2,000 input tokens per turn once history is included, and about 700 visible output tokens back. That is around 2.4M input and 0.84M output tokens a month.

ScenarioBilled tokensMonthly cost
Visible output only2.4M in, 0.84M outabout $9.84
Same, if reasoning triples output2.4M in, 2.5M outabout $19.80
Same volumes on qwen3.7-plus, inside the 256K band2.4M in, 2.5M outabout $4.96

The middle row is the one to sit with, and its multiplier is an assumption rather than a measurement. That is exactly the problem: the amount a reasoning model decides to think is not knowable in advance, so the invoice for a single heavy user on the flagship can exceed the cost of a fixed monthly plan while covering one model family. Metered pricing moves that variance onto you. A flat subscription moves it onto the provider, which for an individual is usually the trade worth making.

Where Qwen genuinely wins

  • The mid tier's price. $0.40 and $1.60 inside the 256K band is a strong rate, and it is where this vendor's affordable reputation is still earned.
  • Flat-rate long context, on the two tiers that actually offer it. qwen3.5-flash at $0.10 and $0.40 across a full million tokens is the quiet bargain on the card.
  • Multimodal input on the flagship, at a rate below several multimodal competitors.
  • Open weights on parts of the line, so a team with GPU capacity and a residency requirement can self-host and none of the pricing above applies. What choosing open weights over a frontier API actually buys covers that decision.
  • The free consumer app. Genuinely capable, genuinely free, no card and nothing to upgrade to.

Where it is not the right default

It is one vendor's model family. Qwen does not give you a different lab's model for the work this one is not suited to, and nothing carries context between them. Qwen against DeepSeek puts the two low-cost open-weight options side by side, which is the comparison most people need before committing to either.

Alibaba Cloud endpoints carry data-residency considerations that matter for regulated industries, client-confidential work, and organisations with a policy on where work product is processed. That is a factual question about your situation rather than a judgement on the model, and self-hosting the open weights is the path that removes it.

The cheaper shape of the same goal

The pattern behind most Qwen pricing searches is someone who wants frontier-quality output without a frontier-sized bill. That instinct was well served by this vendor for most of 2025. In 2026 the flagship is $2.00 and $6.00 with an unpredictable reasoning multiplier on top, and the tiers that are still genuinely cheap are not the flagship.

Perspective AI is $14.99/mo for models from OpenAI, Anthropic, Google, xAI, DeepSeek and Mistral in one app, with mid-conversation switching and memory that carries across them. There is no cloud account to provision, no endpoint region to choose, no reasoning-effort parameter that quietly triples an invoice, and the number is the same every month. For an individual who wanted Qwen because it looked like the affordable route to frontier quality, that is the same goal reached at a fixed price.

The limit on that argument is the same one as everywhere else on this site: if you are running the API at volume rather than chatting, a consumer subscription is a different product and does not replace it. Compare the fixed plans against each other in the per-vendor plan and price coverage.

FAQ

How much does Qwen cost?

The chat app at chat.qwen.ai is free and there is no consumer subscription of any kind. Paid access is metered API usage through Alibaba Cloud Model Studio, where rates run from $0.05 per 1M input tokens on qwen-flash to $2.00 input and $6.00 output on the qwen3.8-max flagship. Read from Alibaba Cloud’s own Model Studio pricing documentation on 2026-08-19.

Is there a Qwen subscription plan like ChatGPT Plus?

No. Alibaba has never shipped a consumer subscription for Qwen, so the product this question is looking for does not exist. You either use the free app with no commitments, or you become an Alibaba Cloud customer with API keys, an endpoint region and a metered invoice. The closest equivalent is a multi-model subscription that includes Qwen.

Is Qwen still cheap compared with the frontier models?

In its mid tier, yes. At the top, not in the way its reputation suggests. qwen3.8-max at $2.00 and $6.00 per 1M tokens is priced in the same band as frontier models rather than well below them, and it is five times the input cost of qwen3.7-plus. There is a 40-fold spread between the cheapest and most expensive input rates inside one product family, so the cheap Qwen and the frontier-class Qwen are different line items.

Does Qwen charge a long-context surcharge?

On the tier most people would pick, yes, and this is widely reported the other way round. qwen3.7-plus is banded: $0.40 and $1.60 up to 256,000 tokens, then $1.20 and $4.80 from 256,000 to a million. That is a threefold step exactly where long-document work lives. Two tiers do hold one rate across the whole window, qwen3.8-max at $2.00 and $6.00 and qwen3.5-flash at $0.10 and $0.40.

Why is my Qwen bill higher than my token estimate?

Usually reasoning output, which bills at the full output rate, so a response that renders as 700 tokens can bill several times that. The reasoning effort is a request parameter, so check what your client sends by default rather than assuming it is low. Two levers pull the other way: batch invocation takes roughly half off input and output where latency permits, and a cache hit costs about 10% of the standard input rate.

Does Qwen have a free tier for developers?

Yes, but it is region-locked. The trial quota, widely reported as a million tokens per eligible model for 90 days, exists only on the International endpoint served from Singapore. Provision in another region and there is no free allowance at all, and it is not transferable afterwards. Region also interacts with data-residency policy, so the right endpoint is often not the one carrying the credit.

Or skip the choice

Frontier models at a fixed price, $14.99/month.

No cloud account, no endpoint region, and no reasoning-effort parameter that quietly triples an invoice. One subscription covering models from OpenAI, Anthropic, Google, xAI, DeepSeek and Mistral, at the same number every month.

Launch app →