qwen

Qwen: Qwen3 VL 32B Instruct

Qwen3 VL 32B Instruct is capable of handling both text and image inputs within a context length of 131,072 tokens and supports tool integration; however, it does not offer reasoning or structured output capabilities. This model excels in processing diverse input modalities but lacks advanced analytical features. For those on a budget who need to process mixed media content without the need for complex reasoning or structured data outputs, Qwen3 VL 32B Instruct is a viable option. Given its blended benchmark score of 16.8 and pricing at $0.104 per million input tokens and $0.416 per million output tokens, it offers a reasonable cost-effectiveness for users prioritizing simplicity and performance in specific use cases.

Quality Score
91/100
price + capability + benchmarks
Input Price
$0.10
per 1M tokens
Output Price
$0.42
per 1M tokens
Context Window
131,072
tokens

Benchmark results

Independent, published benchmarks. Blended score 16.8 across 1 benchmark, last refreshed 2026-07-30. How scoring works →

BenchmarkMeasuresScore
AI Index broad capability composite 18.3
Model ID
qwen/qwen3-vl-32b-instruct
Vendor
qwen
Released
October 2025
Tokenizer
Qwen
Input Modalities
text, image
Output Modalities
text
Max Output
32,768 tokens
Tool Calling
✓ supported
Structured Output
✓ supported
Reasoning Mode
not supported
Vision
✓ accepts images
Audio
no
Moderated
no

What it costs in practice

Computed from the current $0.10/M input and $0.42/M output rates. Run your own numbers →

JobTokensCost
Summarize a 50-page report 30k in / 1.5k out under $0.01
Classify 1,000 customer emails 500k in / 50k out $0.07
A month of a busy support chatbot 5M in / 2M out $1.35

Price & spec history

Tracked daily by PicksByModel since 2026-07-17.

DateInput /MOutput /MContext
2026-07-22 $0.10 $0.42 131,072
2026-07-17 $0.10 $0.42 262,144

Similar models

Quick answers

How much does Qwen: Qwen3 VL 32B Instruct cost?
$0.10 per million input tokens and $0.42 per million output tokens.
What is Qwen: Qwen3 VL 32B Instruct's context window?
131,072 tokens, roughly 196 pages of text in a single request.
Does Qwen: Qwen3 VL 32B Instruct support tool calling?
It supports tool calling, structured output.
Can Qwen: Qwen3 VL 32B Instruct process images?
Yes, it accepts image input.