Zhipu AI (Z.ai)

GLM 5.3 FlashX

Generally available1M contextReasoning

Summary

GLM 5.3 FlashX is a language model from Zhipu AI (Z.ai), released on 18 Sept 2026. It accepts text, images and video and generates text, with a 1M-token context window and up to 131K output tokens per response. As of 4 Oct 2026, the lowest listed API price is $0.37 per million input tokens and $1.25 per million output tokens (OpenRouter), across 1 provider we track.

GLM-5.3-FlashX is the high-speed variant of Z.ai's GLM-5.3-Flash, a native multimodal model delivering inference speeds of up to 200 tokens/s. Built on the same hybrid sparse and linear attention architecture...

Last verified 4 Oct 2026. Every figure on this page links to its source.

Specifications

DeveloperZhipu AI (Z.ai)
Release date18 Sept 2026 [OpenRouter]
StatusGenerally available
Context window1M (1,048,576 tokens) [OpenRouter]
Max output131K tokens [OpenRouter]
InputText, images and video [OpenRouter]
OutputText [OpenRouter]
Reasoning modeYes [OpenRouter]
Tool callingYes [OpenRouter]
Structured outputNo [OpenRouter]

GLM 5.3 FlashX API pricing by provider

USD per million tokens. Sorted by input price.

ProviderInputOutputCached inputContextSource
OpenRouter
z-ai/glm-5.3-flashx
$0.37$1.25$0.091MOpenRouter, 4 Oct 2026

Frequently asked questions

What is the context window of GLM 5.3 FlashX?

GLM 5.3 FlashX has a context window of 1,048,576 tokens (1M) and can generate up to 131,072 tokens in a single response, according to OpenRouter.

How much does the GLM 5.3 FlashX API cost?

GLM 5.3 FlashX is listed from $0.37 per million input tokens and $1.25 per million output tokens (OpenRouter). Prices last verified 4 Oct 2026.

When was GLM 5.3 FlashX released?

Zhipu AI (Z.ai) released GLM 5.3 FlashX on 18 Sept 2026, according to OpenRouter.

Change history

No changes detected since we started tracking this model. We re-check every source daily.

Sources