GLM-4.6V-Flash
Summary
GLM-4.6V-Flash is an open-weight language model from Zhipu AI (Z.ai), released on 8 Dec 2025. It accepts text, images and video and generates text, with a 128K-token context window and up to 33K output tokens per response. As of 4 Oct 2026, the lowest listed API price is $0.30 per million input tokens and $0.90 per million output tokens (Hugging Face Inference Providers), across 2 providers we track.
Lightweight GLM vision model for visual reasoning, documents, and multimodal agents
Last verified 4 Oct 2026. Every figure on this page links to its source.
Specifications
| Developer | Zhipu AI (Z.ai) |
|---|---|
| Release date | 8 Dec 2025 [models.dev] |
| Status | Generally available [models.dev] |
| Weights | Open weights [models.dev] |
| Context window | 128K (128,000 tokens) [models.dev] |
| Max output | 33K tokens [models.dev] |
| Input | Text, images and video [models.dev] |
| Output | Text [models.dev] |
| Reasoning mode | Yes [models.dev] |
| Tool calling | Yes [models.dev] |
GLM-4.6V-Flash API pricing by provider
USD per million tokens. Sorted by input price.
| Provider | Input | Output | Cached input | Context | Source |
|---|---|---|---|---|---|
| Z.ai API glm-4.6v-flash | Free | Free | Free | 128K | models.dev, 4 Oct 2026 |
| Hugging Face Inference Providers zai-org/GLM-4.6V-Flash | $0.30 | $0.90 | – | 131K | models.dev, 4 Oct 2026 |
Frequently asked questions
What is the context window of GLM-4.6V-Flash?
GLM-4.6V-Flash has a context window of 128,000 tokens (128K) and can generate up to 32,768 tokens in a single response, according to models.dev.
How much does the GLM-4.6V-Flash API cost?
Through Z.ai API, GLM-4.6V-Flash costs Free per million input tokens and Free per million output tokens. Prices last verified 4 Oct 2026.
Is GLM-4.6V-Flash open source?
GLM-4.6V-Flash is an open-weight model. Check the license terms before commercial use.
When was GLM-4.6V-Flash released?
Zhipu AI (Z.ai) released GLM-4.6V-Flash on 8 Dec 2025, according to models.dev.
Which providers offer GLM-4.6V-Flash?
We track GLM-4.6V-Flash on 2 providers: Z.ai API, Hugging Face Inference Providers.
Change history
No changes detected since we started tracking this model. We re-check every source daily.