# GLM-4.6V-Flash

> GLM-4.6V-Flash is an open-weight language model from Zhipu AI (Z.ai), released on 8 Dec 2025. It accepts text, images and video and generates text, with a 128K-token context window and up to 33K output tokens per response. As of 4 Oct 2026, the lowest listed API price is $0.30 per million input tokens and $0.90 per million output tokens (Hugging Face Inference Providers), across 2 providers we track.

Source page: https://www.aimodel.directory/models/glm-4-6v-flash
Last verified: 4 Oct 2026

## Specifications

- **Developer:** Zhipu AI (Z.ai)
- **Release date:** 8 Dec 2025 (source: models.dev, https://github.com/sst/models.dev/blob/dev/providers/zai/models/glm-4.6v-flash.toml)
- **Status:** Generally available (source: models.dev, https://github.com/sst/models.dev/blob/dev/providers/zai/models/glm-4.6v-flash.toml)
- **Weights:** Open weights (source: models.dev, https://github.com/sst/models.dev/blob/dev/providers/zai/models/glm-4.6v-flash.toml)
- **Context window:** 128K (128,000 tokens) (source: models.dev, https://github.com/sst/models.dev/blob/dev/providers/zai/models/glm-4.6v-flash.toml)
- **Max output:** 33K tokens (source: models.dev, https://github.com/sst/models.dev/blob/dev/providers/zai/models/glm-4.6v-flash.toml)
- **Input:** Text, images and video (source: models.dev, https://github.com/sst/models.dev/blob/dev/providers/zai/models/glm-4.6v-flash.toml)
- **Output:** Text (source: models.dev, https://github.com/sst/models.dev/blob/dev/providers/zai/models/glm-4.6v-flash.toml)
- **Reasoning mode:** Yes (source: models.dev, https://github.com/sst/models.dev/blob/dev/providers/zai/models/glm-4.6v-flash.toml)
- **Tool calling:** Yes (source: models.dev, https://github.com/sst/models.dev/blob/dev/providers/zai/models/glm-4.6v-flash.toml)

## API pricing (USD per 1M tokens)

| Provider | Input | Output | Cached input | Context |
|---|---|---|---|---|
| Z.ai API | Free | Free | Free | 128K |
| Hugging Face Inference Providers | $0.30 | $0.90 | – | 131K |

## FAQ

### What is the context window of GLM-4.6V-Flash?

GLM-4.6V-Flash has a context window of 128,000 tokens (128K) and can generate up to 32,768 tokens in a single response, according to models.dev.

### How much does the GLM-4.6V-Flash API cost?

Through Z.ai API, GLM-4.6V-Flash costs Free per million input tokens and Free per million output tokens. Prices last verified 4 Oct 2026.

### Is GLM-4.6V-Flash open source?

GLM-4.6V-Flash is an open-weight model. Check the license terms before commercial use.

### When was GLM-4.6V-Flash released?

Zhipu AI (Z.ai) released GLM-4.6V-Flash on 8 Dec 2025, according to models.dev.

### Which providers offer GLM-4.6V-Flash?

We track GLM-4.6V-Flash on 2 providers: Z.ai API, Hugging Face Inference Providers.

---
Data licensed CC BY 4.0. Cite as: AI Model Directory, "GLM-4.6V-Flash", https://www.aimodel.directory/models/glm-4-6v-flash