Models
Enterprise
Subscribe
Resource
Documentation
Console

z.ai/glm-5.3-flash

zhipu

Z.ai: glm-5.3-flash

GLM-5.3-Flash was released by Z.ai in August 2026. The model is the first natively multimodal model in the GLM-5 series, with 320B total parameters, 18B active parameters, hybrid attention, and support for text, image, and video inputs.

Modetext → text
Input price$0.15 per M tokens
Output price$0.5 per M tokens
Context length
Weekly usage0 tokens
Listed at

Z.ai

Latency
Throughput
Upload rate
Context length
Max output
Input price$0.15per 1M tokens
Output price$0.5per 1M tokens
Cache read$0.0045per 1M tokens
Cache write$0per 1M tokens

Performance

Avg TPS
Avg latency
Avg success rate
TPS
TTFT
Latency
Success rate

Success rate

Speed

Code Example

Replace <YOUR_API_KEY> with the API key generated on your token management page.

Authentication

All requests must include the Authorization: Bearer <TOKEN> header.

Endpoint Path

POST

Supported Parameters

ParameterTypeDefaultDescription

Rate Limits

GroupRPMTPMRPD

RPM = requests per minute, TPM = tokens per minute, RPD = requests per day. Limits apply by token group.

Pricing

GroupBilling typePrice summary