GLM-4.7-Flash
Summary
GLM-4.7-Flash is an open-weight language model from Zhipu AI (Z.ai), released on 19 Jan 2026. It accepts text and generates text, with a 200K-token context window and up to 131K output tokens per response. As of 4 Oct 2026, the lowest listed API price is $0.06 per million input tokens and $0.40 per million output tokens (Cloudflare Workers AI), across 7 providers we track.
Budget GLM lane for fast coding help, routing, and everyday automation
Last verified 4 Oct 2026. Every figure on this page links to its source.
Specifications
| Developer | Zhipu AI (Z.ai) |
|---|---|
| Release date | 19 Jan 2026 [models.dev] |
| Status | Generally available [models.dev] |
| Weights | Open weights [Hugging Face] |
| License | mit [Hugging Face] |
| Parameters | 31.2B [Hugging Face] |
| Context window | 200K (200,000 tokens) [models.dev] |
| Max output | 131K tokens [models.dev] |
| Knowledge cutoff | Apr 2025 [models.dev] |
| Input | Text [models.dev] |
| Output | Text [models.dev] |
| Reasoning mode | Yes [models.dev] |
| Tool calling | Yes [models.dev] |
| Hugging Face | zai-org/GLM-4.7-Flash [Hugging Face] |
GLM-4.7-Flash API pricing by provider
USD per million tokens. Sorted by input price.
| Provider | Input | Output | Cached input | Context | Source |
|---|---|---|---|---|---|
| Z.ai API glm-4.7-flash | Free | Free | Free | 200K | models.dev, 4 Oct 2026 |
| Hugging Face Inference Providers zai-org/GLM-4.7-Flash | Free | Free | – | 200K | models.dev, 4 Oct 2026 |
| Cloudflare Workers AI @cf/zai-org/glm-4.7-flash | $0.06 | $0.40 | – | 131K | models.dev, 4 Oct 2026 |
| OpenRouter z-ai/glm-4.7-flash | $0.06 | $0.40 | – | 131K | OpenRouter, 4 Oct 2026 |
| Amazon Bedrock zai.glm-4.7-flash | $0.07 | $0.40 | – | 203K | models.dev, 4 Oct 2026 |
| Novita AI zai-org/glm-4.7-flash | $0.07 | $0.40 | $0.01 | 200K | models.dev, 4 Oct 2026 |
| Vercel AI Gateway zai/glm-4.7-flash | $0.07 | $0.40 | – | 200K | models.dev, 4 Oct 2026 |
Frequently asked questions
What is the context window of GLM-4.7-Flash?
GLM-4.7-Flash has a context window of 200,000 tokens (200K) and can generate up to 131,072 tokens in a single response, according to models.dev.
How much does the GLM-4.7-Flash API cost?
Through Z.ai API, GLM-4.7-Flash costs Free per million input tokens and Free per million output tokens. Prices last verified 4 Oct 2026.
Is GLM-4.7-Flash open source?
GLM-4.7-Flash is an open-weight model: its weights are published on Hugging Face as zai-org/GLM-4.7-Flash under the mit license. Check the license terms before commercial use.
When was GLM-4.7-Flash released?
Zhipu AI (Z.ai) released GLM-4.7-Flash on 19 Jan 2026, according to models.dev.
Which providers offer GLM-4.7-Flash?
We track GLM-4.7-Flash on 7 providers: Z.ai API, Hugging Face Inference Providers, Cloudflare Workers AI, OpenRouter, Amazon Bedrock, Novita AI, Vercel AI Gateway.
Change history
No changes detected since we started tracking this model. We re-check every source daily.
Sources
- https://openrouter.ai/z-ai/glm-4.7-flash
- https://github.com/sst/models.dev/blob/dev/providers/zai/models/glm-4.7-flash.toml
- https://huggingface.co/zai-org/GLM-4.7-Flash
- https://github.com/sst/models.dev/blob/dev/providers/huggingface/models/zai-org/GLM-4.7-Flash.toml
- https://github.com/sst/models.dev/blob/dev/providers/cloudflare-workers-ai/models/@cf/zai-org/glm-4.7-flash.toml
- https://github.com/sst/models.dev/blob/dev/providers/amazon-bedrock/models/zai.glm-4.7-flash.toml
- https://github.com/sst/models.dev/blob/dev/providers/novita-ai/models/zai-org/glm-4.7-flash.toml
- https://github.com/sst/models.dev/blob/dev/providers/vercel/models/zai/glm-4.7-flash.toml