~z-ai

Z.ai: GLM Flash Latest

The GLM Flash Latest model is offered by vendor ~z-ai and supports text, image, and video input modalities. It has a context length of 1310720 tokens and can generate up to 943718 completion tokens. The model enables the use of tools and reasoning capabilities. For users seeking a multi-modal AI solution that supports various input types, this model may be worth shortlisting due to its unique feature set. However, its pricing point, with $0.075/M input and $0.25/M output tokens, will also play a significant role in the decision-making process. As no independent benchmark coverage is available, further evaluation is necessary to fully assess its performance.

Quality Score
100/100
price + capability + benchmarks
Input Price
$0.07
per 1M tokens · checked 2026-09-04
Output Price
$0.25
per 1M tokens · checked 2026-09-04
Context Window
1,310,720
tokens
Model ID
~z-ai/glm-flash-latest
Vendor
~z-ai
Released
August 2026
Tokenizer
Router
Input Modalities
text, image, video
Output Modalities
text
Max Output
943,718 tokens
Tool Calling
✓ supported
Structured Output
✓ supported
Reasoning Mode
✓ supported
Vision
✓ accepts images
Audio
no
Moderated
no

What it costs in practice

Computed from the current $0.07/M input and $0.25/M output rates. Run your own numbers →

JobTokensCost
Summarize a 50-page report 30k in / 1.5k out under $0.01
Classify 1,000 customer emails 500k in / 50k out $0.05
A month of a busy support chatbot 5M in / 2M out $0.88

Strong choice for

Category rankings

Where Z.ai: GLM Flash Latest places across the 6 categories it ranks in. How we rank →

#CategoryScore
#2 Real-Time ChatLatency · of 25 ranked 118
#4 Video Auto-TaggingVideo · of 25 ranked 123
#9 Self-Hosted / LocalCost · of 25 ranked 117
#13 Social Media PostsWriting · of 25 ranked 119
#13 Voice Assistant BackendVoice · of 25 ranked 123
#13 Cheap Bulk InferenceCost · of 25 ranked 137

Similar models

Quick answers

How much does Z.ai: GLM Flash Latest cost?
$0.07 per million input tokens and $0.25 per million output tokens.
What is Z.ai: GLM Flash Latest's context window?
1,310,720 tokens, roughly 1,966 pages of text in a single request.
Does Z.ai: GLM Flash Latest support tool calling?
It supports tool calling, structured output, a reasoning mode.
Can Z.ai: GLM Flash Latest process images?
Yes, it accepts image input.