qwen

Qwen: Qwen2.5 VL 72B Instruct

Qwen2.5 VL 72B Instruct is a large language model capable of handling text and image inputs within a context length of up to 128,000 tokens; however, it does not support tools, reasoning, or structured output. While its extensive input modalities make it versatile for various tasks involving both text and images, the lack of tool integration and reasoning abilities limits its utility in scenarios requiring interaction with external data sources or complex logical processing. Given its pricing at $0.8 per million input tokens and $1.0 per million output tokens, Qwen2.5 VL 72B Instruct is more suitable for users who need a robust text and image handling model within their budget but do not require benchmark-tested performance or advanced reasoning capabilities. Its unproven benchmark standing highlights that its effectiveness in specific tasks remains to be independently validated.

Quality Score
80/100
price + capability + benchmarks
Input Price
$0.80
per 1M tokens
Output Price
$1.00
per 1M tokens
Context Window
128,000
tokens
Model ID
qwen/qwen2.5-vl-72b-instruct
Vendor
qwen
Released
February 2025
Tokenizer
Qwen
Input Modalities
text, image
Output Modalities
text
Max Output
128,000 tokens
Tool Calling
not supported
Structured Output
✓ supported
Reasoning Mode
not supported
Vision
✓ accepts images
Audio
no
Moderated
no

What it costs in practice

Computed from the current $0.80/M input and $1.00/M output rates. Run your own numbers →

JobTokensCost
Summarize a 50-page report 30k in / 1.5k out $0.03
Classify 1,000 customer emails 500k in / 50k out $0.45
A month of a busy support chatbot 5M in / 2M out $6.00

Price & spec history

Tracked daily by PicksByModel since 2026-07-17.

DateInput /MOutput /MContext
2026-07-22 $0.80 $1.00 128,000
2026-07-17 $0.80 $1.00 131,072

Similar models

Quick answers

How much does Qwen: Qwen2.5 VL 72B Instruct cost?
$0.80 per million input tokens and $1.00 per million output tokens.
What is Qwen: Qwen2.5 VL 72B Instruct's context window?
128,000 tokens, roughly 192 pages of text in a single request.
Does Qwen: Qwen2.5 VL 72B Instruct support tool calling?
It supports structured output.
Can Qwen: Qwen2.5 VL 72B Instruct process images?
Yes, it accepts image input.