z-ai

Z.ai: GLM 5V Turbo

The Z.ai GLM 5V Turbo is a multi-modal AI model that can handle text, image, and video inputs up to a context length of 202752 tokens. It supports tools and reasoning capabilities. This model may be worth shortlisting for those who prioritize its unique multi-modal capabilities and don't require structured output. However, with a price point of $1.2 per million input tokens and $4.0 per million output tokens, it's essential to evaluate whether the costs align with your specific use case. Despite scoring 57.6 on blended benchmark tests across one independent measure, its overall performance is uncertain due to limited coverage data.

Quality Score
100/100
price + capability + benchmarks
Input Price
$1.20
per 1M tokens · checked 2026-09-04
Output Price
$4.00
per 1M tokens · checked 2026-09-04
Context Window
202,752
tokens

Benchmark results

Independent, published benchmarks. Blended score 57.6 across 1 benchmark, last refreshed 2026-09-04. How scoring works →

BenchmarkMeasuresScore
AI Index broad capability composite 58.3
Model ID
z-ai/glm-5v-turbo
Vendor
z-ai
Released
April 2026
Tokenizer
Other
Input Modalities
image, text, video
Output Modalities
text
Max Output
131,072 tokens
Tool Calling
✓ supported
Structured Output
✓ supported
Reasoning Mode
✓ supported
Vision
✓ accepts images
Audio
no
Moderated
no

What it costs in practice

Computed from the current $1.20/M input and $4.00/M output rates. Run your own numbers →

JobTokensCost
Summarize a 50-page report 30k in / 1.5k out $0.04
Classify 1,000 customer emails 500k in / 50k out $0.80
A month of a busy support chatbot 5M in / 2M out $14.00

Similar models

Quick answers

How much does Z.ai: GLM 5V Turbo cost?
$1.20 per million input tokens and $4.00 per million output tokens.
What is Z.ai: GLM 5V Turbo's context window?
202,752 tokens, roughly 304 pages of text in a single request.
Does Z.ai: GLM 5V Turbo support tool calling?
It supports tool calling, structured output, a reasoning mode.
Can Z.ai: GLM 5V Turbo process images?
Yes, it accepts image input.