Zhipu AI (Z.ai)

GLM-5.3-Flash

Generally availableOpen weights1M contextReasoning

Summary

GLM-5.3-Flash is an open-weight language model from Zhipu AI (Z.ai), released on 26 Aug 2026. It accepts text, images, video and PDFs and generates text, with a 1M-token context window and up to 131K output tokens per response. As of 4 Oct 2026, the lowest listed API price is $0.15 per million input tokens and $0.50 per million output tokens (Z.ai API), across 13 providers we track.

Native multimodal GLM model for efficient coding and long-horizon agent tasks

Last verified 4 Oct 2026. Every figure on this page links to its source.

Specifications

DeveloperZhipu AI (Z.ai)
Release date26 Aug 2026 [models.dev]
StatusGenerally available [models.dev]
WeightsOpen weights [Hugging Face]
Licensemit [Hugging Face]
Parameters321B [Hugging Face]
Context window1M (1,000,000 tokens) [models.dev]
Max output131K tokens [models.dev]
InputText, images, video and PDFs [models.dev]
OutputText [models.dev]
Reasoning modeYes [models.dev]
Tool callingYes [models.dev]
Structured outputYes [models.dev]
Hugging Facezai-org/GLM-5.3-Flash [Hugging Face]

GLM-5.3-Flash API pricing by provider

USD per million tokens. Sorted by input price.

ProviderInputOutputCached inputContextSource
NVIDIA NIM
z-ai/glm-5.3-flash
FreeFree–1Mmodels.dev, 4 Oct 2026
Z.ai API
glm-5.3-flash
$0.15$0.50$0.031Mmodels.dev, 4 Oct 2026
Cloudflare Workers AI
@cf/zai-org/glm-5.3-flash
$0.15$0.50$0.031Mmodels.dev, 4 Oct 2026
SiliconFlow
zai-org/GLM-5.3-Flash
$0.15$0.50$0.031Mmodels.dev, 4 Oct 2026
Hugging Face Inference Providers
zai-org/GLM-5.3-Flash
$0.15$0.50–1Mmodels.dev, 4 Oct 2026
Nebius AI Studio
zai-org/GLM-5.3-Flash
$0.15$0.50$0.151Mmodels.dev, 4 Oct 2026
Together AI
zai-org/GLM-5.3-Flash
$0.15$0.50$0.031Mmodels.dev, 4 Oct 2026
Fireworks AI
accounts/fireworks/models/glm-5p3-flash
$0.15$0.50$0.031Mmodels.dev, 4 Oct 2026
FriendliAI
zai-org/GLM-5.3-Flash
$0.15$0.50$0.031Mmodels.dev, 4 Oct 2026
Fireworks AI
accounts/fireworks/routers/glm-flash-latest
$0.15$0.50$0.031Mmodels.dev, 4 Oct 2026
DeepInfra
zai-org/GLM-5.3-Flash
$0.15$0.50$0.031Mmodels.dev, 4 Oct 2026
Baseten
zai-org/GLM-5.3-Flash
$0.15$0.50–1Mmodels.dev, 4 Oct 2026
OpenRouter
z-ai/glm-5.3-flash
$0.15$0.50$0.031MOpenRouter, 4 Oct 2026
Vercel AI Gateway
zai/glm-5.3-flash
$0.15$0.50$0.031Mmodels.dev, 4 Oct 2026
Z.ai API
glm-5.3-flashx
$0.37$1.25$0.071Mmodels.dev, 4 Oct 2026
Vercel AI Gateway
zai/glm-5.3-flashx
$0.37$1.25$0.071Mmodels.dev, 4 Oct 2026

Frequently asked questions

What is the context window of GLM-5.3-Flash?

GLM-5.3-Flash has a context window of 1,000,000 tokens (1M) and can generate up to 131,072 tokens in a single response, according to models.dev.

How much does the GLM-5.3-Flash API cost?

Through Z.ai API, GLM-5.3-Flash costs $0.15 per million input tokens and $0.50 per million output tokens. Prices last verified 4 Oct 2026.

Is GLM-5.3-Flash open source?

GLM-5.3-Flash is an open-weight model: its weights are published on Hugging Face as zai-org/GLM-5.3-Flash under the mit license. Check the license terms before commercial use.

When was GLM-5.3-Flash released?

Zhipu AI (Z.ai) released GLM-5.3-Flash on 26 Aug 2026, according to models.dev.

Which providers offer GLM-5.3-Flash?

We track GLM-5.3-Flash on 13 providers: NVIDIA NIM, Z.ai API, Cloudflare Workers AI, SiliconFlow, Hugging Face Inference Providers, Nebius AI Studio, Together AI, Fireworks AI, FriendliAI, DeepInfra, Baseten, OpenRouter, Vercel AI Gateway.

Change history

No changes detected since we started tracking this model. We re-check every source daily.

Sources