Zhipu AI (Z.ai)

GLM-4.5-Flash

Generally availableOpen weights131K contextReasoning

Summary

GLM-4.5-Flash is an open-weight language model from Zhipu AI (Z.ai), released on 28 Jul 2025. It accepts text and generates text, with a 131K-token context window and up to 98K output tokens per response.

Efficient GLM model for fast reasoning, coding, and agent workflows

Last verified 4 Oct 2026. Every figure on this page links to its source.

Specifications

DeveloperZhipu AI (Z.ai)
Release date28 Jul 2025 [models.dev]
StatusGenerally available [models.dev]
WeightsOpen weights [models.dev]
Context window131K (131,072 tokens) [models.dev]
Max output98K tokens [models.dev]
Knowledge cutoffApr 2025 [models.dev]
InputText [models.dev]
OutputText [models.dev]
Reasoning modeYes [models.dev]
Tool callingYes [models.dev]

GLM-4.5-Flash API pricing by provider

USD per million tokens. Sorted by input price.

ProviderInputOutputCached inputContextSource
Z.ai API
glm-4.5-flash
FreeFreeFree131Kmodels.dev, 4 Oct 2026

Frequently asked questions

What is the context window of GLM-4.5-Flash?

GLM-4.5-Flash has a context window of 131,072 tokens (131K) and can generate up to 98,304 tokens in a single response, according to models.dev.

Is GLM-4.5-Flash open source?

GLM-4.5-Flash is an open-weight model. Check the license terms before commercial use.

When was GLM-4.5-Flash released?

Zhipu AI (Z.ai) released GLM-4.5-Flash on 28 Jul 2025, according to models.dev.

Change history

No changes detected since we started tracking this model. We re-check every source daily.

Sources