Z.ai: GLM Flash Latest
The GLM Flash Latest model is offered by vendor ~z-ai and supports text, image, and video input modalities. It has a context length of 1310720 tokens and can generate up to 943718 completion tokens. The model enables the use of tools and reasoning capabilities. For users seeking a multi-modal AI solution that supports various input types, this model may be worth shortlisting due to its unique feature set. However, its pricing point, with $0.075/M input and $0.25/M output tokens, will also play a significant role in the decision-making process. As no independent benchmark coverage is available, further evaluation is necessary to fully assess its performance.
- Model ID
- ~z-ai/glm-flash-latest
- Vendor
- ~z-ai
- Released
- August 2026
- Tokenizer
- Router
- Input Modalities
- text, image, video
- Output Modalities
- text
- Max Output
- 943,718 tokens
- Tool Calling
- ✓ supported
- Structured Output
- ✓ supported
- Reasoning Mode
- ✓ supported
- Vision
- ✓ accepts images
- Audio
- no
- Moderated
- no
What it costs in practice
Computed from the current $0.07/M input and $0.25/M output rates. Run your own numbers →
| Job | Tokens | Cost |
|---|---|---|
| Summarize a 50-page report | 30k in / 1.5k out | under $0.01 |
| Classify 1,000 customer emails | 500k in / 50k out | $0.05 |
| A month of a busy support chatbot | 5M in / 2M out | $0.88 |
Strong choice for
Category rankings
Where Z.ai: GLM Flash Latest places across the 6 categories it ranks in. How we rank →
| # | Category | Score |
|---|---|---|
| #2 | Real-Time ChatLatency · of 25 ranked | 118 |
| #4 | Video Auto-TaggingVideo · of 25 ranked | 123 |
| #9 | Self-Hosted / LocalCost · of 25 ranked | 117 |
| #13 | Social Media PostsWriting · of 25 ranked | 119 |
| #13 | Voice Assistant BackendVoice · of 25 ranked | 123 |
| #13 | Cheap Bulk InferenceCost · of 25 ranked | 137 |
Similar models
Quick answers
- How much does Z.ai: GLM Flash Latest cost?
- $0.07 per million input tokens and $0.25 per million output tokens.
- What is Z.ai: GLM Flash Latest's context window?
- 1,310,720 tokens, roughly 1,966 pages of text in a single request.
- Does Z.ai: GLM Flash Latest support tool calling?
- It supports tool calling, structured output, a reasoning mode.
- Can Z.ai: GLM Flash Latest process images?
- Yes, it accepts image input.