z-ai

Z.ai: GLM 4.6V

The Z.ai: GLM 4.6V model is a versatile AI tool that can process multiple input types, including images, text, and video. It operates with a context length of 131072 tokens and supports reasoning capabilities. When evaluating this option, consider its pricing structure, which charges $0.3 per million input tokens and $0.9 per million output tokens. The model's blended benchmark score of 9.3 across one benchmark suggests it performs reasonably well in certain tasks. Those seeking a tool with broad modality support and reasoning capabilities might find Z.ai: GLM 4.6V worth shortlisting, particularly given its competitive pricing.

Quality Score
100/100
price + capability + benchmarks
Input Price
$0.30
per 1M tokens · checked 2026-09-08
Output Price
$0.90
per 1M tokens · checked 2026-09-08
Context Window
131,072
tokens

Benchmark results

Independent, published benchmarks. Blended score 9.3 across 1 benchmark, last refreshed 2026-09-08. How scoring works →

BenchmarkMeasuresScore
AI Index broad capability composite 13.8
Model ID
z-ai/glm-4.6v
Vendor
z-ai
Released
December 2025
Tokenizer
Other
Input Modalities
image, text, video
Output Modalities
text
Max Output
32,768 tokens
Tool Calling
✓ supported
Structured Output
✓ supported
Reasoning Mode
✓ supported
Vision
✓ accepts images
Audio
no
Moderated
no

What it costs in practice

Computed from the current $0.30/M input and $0.90/M output rates. Run your own numbers →

JobTokensCost
Summarize a 50-page report 30k in / 1.5k out $0.01
Classify 1,000 customer emails 500k in / 50k out $0.20
A month of a busy support chatbot 5M in / 2M out $3.30

Similar models

Quick answers

How much does Z.ai: GLM 4.6V cost?
$0.30 per million input tokens and $0.90 per million output tokens.
What is Z.ai: GLM 4.6V's context window?
131,072 tokens, roughly 196 pages of text in a single request.
Does Z.ai: GLM 4.6V support tool calling?
It supports tool calling, structured output, a reasoning mode.
Can Z.ai: GLM 4.6V process images?
Yes, it accepts image input.