# GLM-5.3-Flash

> GLM-5.3-Flash is an open-weight language model from Zhipu AI (Z.ai), released on 26 Aug 2026. It accepts text, images, video and PDFs and generates text, with a 1M-token context window and up to 131K output tokens per response. As of 4 Oct 2026, the lowest listed API price is $0.15 per million input tokens and $0.50 per million output tokens (Z.ai API), across 13 providers we track.

Source page: https://www.aimodel.directory/models/glm-5-3-flash
Last verified: 4 Oct 2026

## Specifications

- **Developer:** Zhipu AI (Z.ai)
- **Release date:** 26 Aug 2026 (source: models.dev, https://github.com/sst/models.dev/blob/dev/providers/zai/models/glm-5.3-flash.toml)
- **Status:** Generally available (source: models.dev, https://github.com/sst/models.dev/blob/dev/providers/zai/models/glm-5.3-flash.toml)
- **Weights:** Open weights (source: Hugging Face, https://huggingface.co/zai-org/GLM-5.3-Flash)
- **License:** mit (source: Hugging Face, https://huggingface.co/zai-org/GLM-5.3-Flash)
- **Parameters:** 321B (source: Hugging Face, https://huggingface.co/zai-org/GLM-5.3-Flash)
- **Context window:** 1M (1,000,000 tokens) (source: models.dev, https://github.com/sst/models.dev/blob/dev/providers/zai/models/glm-5.3-flash.toml)
- **Max output:** 131K tokens (source: models.dev, https://github.com/sst/models.dev/blob/dev/providers/zai/models/glm-5.3-flash.toml)
- **Input:** Text, images, video and PDFs (source: models.dev, https://github.com/sst/models.dev/blob/dev/providers/zai/models/glm-5.3-flash.toml)
- **Output:** Text (source: models.dev, https://github.com/sst/models.dev/blob/dev/providers/zai/models/glm-5.3-flash.toml)
- **Reasoning mode:** Yes (source: models.dev, https://github.com/sst/models.dev/blob/dev/providers/zai/models/glm-5.3-flash.toml)
- **Tool calling:** Yes (source: models.dev, https://github.com/sst/models.dev/blob/dev/providers/zai/models/glm-5.3-flash.toml)
- **Structured output:** Yes (source: models.dev, https://github.com/sst/models.dev/blob/dev/providers/zai/models/glm-5.3-flash.toml)
- **Hugging Face:** zai-org/GLM-5.3-Flash (source: Hugging Face, https://huggingface.co/zai-org/GLM-5.3-Flash)

## API pricing (USD per 1M tokens)

| Provider | Input | Output | Cached input | Context |
|---|---|---|---|---|
| NVIDIA NIM | Free | Free | – | 1M |
| Z.ai API | $0.15 | $0.50 | $0.03 | 1M |
| Cloudflare Workers AI | $0.15 | $0.50 | $0.03 | 1M |
| SiliconFlow | $0.15 | $0.50 | $0.03 | 1M |
| Hugging Face Inference Providers | $0.15 | $0.50 | – | 1M |
| Nebius AI Studio | $0.15 | $0.50 | $0.15 | 1M |
| Together AI | $0.15 | $0.50 | $0.03 | 1M |
| Fireworks AI | $0.15 | $0.50 | $0.03 | 1M |
| FriendliAI | $0.15 | $0.50 | $0.03 | 1M |
| Fireworks AI | $0.15 | $0.50 | $0.03 | 1M |
| DeepInfra | $0.15 | $0.50 | $0.03 | 1M |
| Baseten | $0.15 | $0.50 | – | 1M |
| OpenRouter | $0.15 | $0.50 | $0.03 | 1M |
| Vercel AI Gateway | $0.15 | $0.50 | $0.03 | 1M |
| Z.ai API | $0.37 | $1.25 | $0.07 | 1M |
| Vercel AI Gateway | $0.37 | $1.25 | $0.07 | 1M |

## FAQ

### What is the context window of GLM-5.3-Flash?

GLM-5.3-Flash has a context window of 1,000,000 tokens (1M) and can generate up to 131,072 tokens in a single response, according to models.dev.

### How much does the GLM-5.3-Flash API cost?

Through Z.ai API, GLM-5.3-Flash costs $0.15 per million input tokens and $0.50 per million output tokens. Prices last verified 4 Oct 2026.

### Is GLM-5.3-Flash open source?

GLM-5.3-Flash is an open-weight model: its weights are published on Hugging Face as zai-org/GLM-5.3-Flash under the mit license. Check the license terms before commercial use.

### When was GLM-5.3-Flash released?

Zhipu AI (Z.ai) released GLM-5.3-Flash on 26 Aug 2026, according to models.dev.

### Which providers offer GLM-5.3-Flash?

We track GLM-5.3-Flash on 13 providers: NVIDIA NIM, Z.ai API, Cloudflare Workers AI, SiliconFlow, Hugging Face Inference Providers, Nebius AI Studio, Together AI, Fireworks AI, FriendliAI, DeepInfra, Baseten, OpenRouter, Vercel AI Gateway.

---
Data licensed CC BY 4.0. Cite as: AI Model Directory, "GLM-5.3-Flash", https://www.aimodel.directory/models/glm-5-3-flash