Alibaba (Qwen)

Qwen Flash

Generally availableProprietary1M contextReasoning

Summary

Qwen Flash is a proprietary language model from Alibaba (Qwen), released on 28 Jul 2025. It accepts text and generates text, with a 1M-token context window and up to 33K output tokens per response. As of 4 Oct 2026, the lowest listed API price is $0.05 per million input tokens and $0.40 per million output tokens (Alibaba Cloud Model Studio), across 1 provider we track.

Efficient Qwen model for fast chat, extraction, and high-volume workloads

Last verified 4 Oct 2026. Every figure on this page links to its source.

Specifications

DeveloperAlibaba (Qwen)
Release date28 Jul 2025 [models.dev]
StatusGenerally available [models.dev]
WeightsProprietary (API only) [models.dev]
Context window1M (1,000,000 tokens) [models.dev]
Max output33K tokens [models.dev]
Knowledge cutoffApr 2024 [models.dev]
InputText [models.dev]
OutputText [models.dev]
Reasoning modeYes [models.dev]
Tool callingYes [models.dev]

Qwen Flash API pricing by provider

USD per million tokens. Sorted by input price.

ProviderInputOutputCached inputContextSource
Alibaba Cloud Model Studio
qwen-flash
$0.05$0.40–1Mmodels.dev, 4 Oct 2026

Frequently asked questions

What is the context window of Qwen Flash?

Qwen Flash has a context window of 1,000,000 tokens (1M) and can generate up to 32,768 tokens in a single response, according to models.dev.

How much does the Qwen Flash API cost?

Through Alibaba Cloud Model Studio, Qwen Flash costs $0.05 per million input tokens and $0.40 per million output tokens. Prices last verified 4 Oct 2026.

Is Qwen Flash open source?

No. Qwen Flash is a proprietary model; Alibaba (Qwen) has not released its weights. It is available through APIs including Alibaba Cloud Model Studio.

When was Qwen Flash released?

Alibaba (Qwen) released Qwen Flash on 28 Jul 2025, according to models.dev.

Change history

No changes detected since we started tracking this model. We re-check every source daily.

Sources