Zhipu AI (Z.ai)

GLM-4.7-Flash

Generally availableOpen weights200K contextReasoning

Summary

GLM-4.7-Flash is an open-weight language model from Zhipu AI (Z.ai), released on 19 Jan 2026. It accepts text and generates text, with a 200K-token context window and up to 131K output tokens per response. As of 4 Oct 2026, the lowest listed API price is $0.06 per million input tokens and $0.40 per million output tokens (Cloudflare Workers AI), across 7 providers we track.

Budget GLM lane for fast coding help, routing, and everyday automation

Last verified 4 Oct 2026. Every figure on this page links to its source.

Specifications

DeveloperZhipu AI (Z.ai)
Release date19 Jan 2026 [models.dev]
StatusGenerally available [models.dev]
WeightsOpen weights [Hugging Face]
Licensemit [Hugging Face]
Parameters31.2B [Hugging Face]
Context window200K (200,000 tokens) [models.dev]
Max output131K tokens [models.dev]
Knowledge cutoffApr 2025 [models.dev]
InputText [models.dev]
OutputText [models.dev]
Reasoning modeYes [models.dev]
Tool callingYes [models.dev]
Hugging Facezai-org/GLM-4.7-Flash [Hugging Face]

GLM-4.7-Flash API pricing by provider

USD per million tokens. Sorted by input price.

ProviderInputOutputCached inputContextSource
Z.ai API
glm-4.7-flash
FreeFreeFree200Kmodels.dev, 4 Oct 2026
Hugging Face Inference Providers
zai-org/GLM-4.7-Flash
FreeFree–200Kmodels.dev, 4 Oct 2026
Cloudflare Workers AI
@cf/zai-org/glm-4.7-flash
$0.06$0.40–131Kmodels.dev, 4 Oct 2026
OpenRouter
z-ai/glm-4.7-flash
$0.06$0.40–131KOpenRouter, 4 Oct 2026
Amazon Bedrock
zai.glm-4.7-flash
$0.07$0.40–203Kmodels.dev, 4 Oct 2026
Novita AI
zai-org/glm-4.7-flash
$0.07$0.40$0.01200Kmodels.dev, 4 Oct 2026
Vercel AI Gateway
zai/glm-4.7-flash
$0.07$0.40–200Kmodels.dev, 4 Oct 2026

Frequently asked questions

What is the context window of GLM-4.7-Flash?

GLM-4.7-Flash has a context window of 200,000 tokens (200K) and can generate up to 131,072 tokens in a single response, according to models.dev.

How much does the GLM-4.7-Flash API cost?

Through Z.ai API, GLM-4.7-Flash costs Free per million input tokens and Free per million output tokens. Prices last verified 4 Oct 2026.

Is GLM-4.7-Flash open source?

GLM-4.7-Flash is an open-weight model: its weights are published on Hugging Face as zai-org/GLM-4.7-Flash under the mit license. Check the license terms before commercial use.

When was GLM-4.7-Flash released?

Zhipu AI (Z.ai) released GLM-4.7-Flash on 19 Jan 2026, according to models.dev.

Which providers offer GLM-4.7-Flash?

We track GLM-4.7-Flash on 7 providers: Z.ai API, Hugging Face Inference Providers, Cloudflare Workers AI, OpenRouter, Amazon Bedrock, Novita AI, Vercel AI Gateway.

Change history

No changes detected since we started tracking this model. We re-check every source daily.

Sources