# Qwen3 32B

> Qwen3 32B is an open-weight language model from Alibaba (Qwen), released on Apr 2025. It accepts text and generates text, with a 131K-token context window and up to 16K output tokens per response. As of 4 Oct 2026, the lowest listed API price is $0.08 per million input tokens and $0.28 per million output tokens (DeepInfra), across 7 providers we track.

Source page: https://www.aimodel.directory/models/qwen3-32b
Last verified: 4 Oct 2026

## Specifications

- **Developer:** Alibaba (Qwen)
- **Release date:** Apr 2025 (source: models.dev, https://github.com/sst/models.dev/blob/dev/providers/alibaba/models/qwen3-32b.toml)
- **Status:** Generally available (source: models.dev, https://github.com/sst/models.dev/blob/dev/providers/alibaba/models/qwen3-32b.toml)
- **Weights:** Open weights (source: Hugging Face, https://huggingface.co/Qwen/Qwen3-32B)
- **License:** apache-2.0 (source: Hugging Face, https://huggingface.co/Qwen/Qwen3-32B)
- **Parameters:** 32.8B (source: Hugging Face, https://huggingface.co/Qwen/Qwen3-32B)
- **Context window:** 131K (131,072 tokens) (source: models.dev, https://github.com/sst/models.dev/blob/dev/providers/alibaba/models/qwen3-32b.toml)
- **Max output:** 16K tokens (source: models.dev, https://github.com/sst/models.dev/blob/dev/providers/alibaba/models/qwen3-32b.toml)
- **Knowledge cutoff:** Apr 2025 (source: models.dev, https://github.com/sst/models.dev/blob/dev/providers/alibaba/models/qwen3-32b.toml)
- **Input:** Text (source: models.dev, https://github.com/sst/models.dev/blob/dev/providers/alibaba/models/qwen3-32b.toml)
- **Output:** Text (source: models.dev, https://github.com/sst/models.dev/blob/dev/providers/alibaba/models/qwen3-32b.toml)
- **Reasoning mode:** Yes (source: models.dev, https://github.com/sst/models.dev/blob/dev/providers/alibaba/models/qwen3-32b.toml)
- **Tool calling:** Yes (source: models.dev, https://github.com/sst/models.dev/blob/dev/providers/alibaba/models/qwen3-32b.toml)
- **Hugging Face:** Qwen/Qwen3-32B (source: Hugging Face, https://huggingface.co/Qwen/Qwen3-32B)

## API pricing (USD per 1M tokens)

| Provider | Input | Output | Cached input | Context |
|---|---|---|---|---|
| DeepInfra | $0.08 | $0.28 | – | 41K |
| OpenRouter | $0.08 | $0.28 | – | 41K |
| SiliconFlow | $0.14 | $0.57 | – | 131K |
| Amazon Bedrock | $0.15 | $0.60 | – | 33K |
| Vercel AI Gateway | $0.16 | $0.64 | – | 128K |
| Hugging Face Inference Providers | $0.29 | $0.59 | – | 131K |
| Alibaba Cloud Model Studio | $0.70 | $2.80 | – | 131K |

## FAQ

### What is the context window of Qwen3 32B?

Qwen3 32B has a context window of 131,072 tokens (131K) and can generate up to 16,384 tokens in a single response, according to models.dev.

### How much does the Qwen3 32B API cost?

Through Alibaba Cloud Model Studio, Qwen3 32B costs $0.70 per million input tokens and $2.80 per million output tokens. The lowest listed price is $0.08 input / $0.28 output through DeepInfra. Prices last verified 4 Oct 2026.

### Is Qwen3 32B open source?

Qwen3 32B is an open-weight model: its weights are published on Hugging Face as Qwen/Qwen3-32B under the apache-2.0 license. Check the license terms before commercial use.

### When was Qwen3 32B released?

Alibaba (Qwen) released Qwen3 32B on Apr 2025, according to models.dev.

### Which providers offer Qwen3 32B?

We track Qwen3 32B on 7 providers: DeepInfra, OpenRouter, SiliconFlow, Amazon Bedrock, Vercel AI Gateway, Hugging Face Inference Providers, Alibaba Cloud Model Studio.

---
Data licensed CC BY 4.0. Cite as: AI Model Directory, "Qwen3 32B", https://www.aimodel.directory/models/qwen3-32b