qwen

Qwen: Qwen3 VL 8B Instruct

Qwen3 VL 8B Instruct is designed to process both text and image inputs within a context length of 256,000 tokens; it supports tools for enhanced functionality but lacks reasoning capabilities or structured output. At $0.117 per million input tokens and $0.455 per million output tokens, this model is suitable for those needing versatile modalities and tool integration within a budget. Given its lack of benchmark coverage, users should consider Qwen3 VL 8B Instruct when prioritizing flexibility and cost efficiency over performance metrics; the current pricing makes it accessible for projects that require text and image handling with the support of external tools.

Quality Score
99/100
price + capability + benchmarks
Input Price
$0.12
per 1M tokens
Output Price
$0.46
per 1M tokens
Context Window
256,000
tokens
Model ID
qwen/qwen3-vl-8b-instruct
Vendor
qwen
Released
October 2025
Tokenizer
Qwen3
Input Modalities
image, text
Output Modalities
text
Max Output
32,768 tokens
Tool Calling
✓ supported
Structured Output
✓ supported
Reasoning Mode
not supported
Vision
✓ accepts images
Audio
no
Moderated
no

What it costs in practice

Computed from the current $0.12/M input and $0.46/M output rates. Run your own numbers →

JobTokensCost
Summarize a 50-page report 30k in / 1.5k out under $0.01
Classify 1,000 customer emails 500k in / 50k out $0.08
A month of a busy support chatbot 5M in / 2M out $1.50

Similar models

Quick answers

How much does Qwen: Qwen3 VL 8B Instruct cost?
$0.12 per million input tokens and $0.46 per million output tokens.
What is Qwen: Qwen3 VL 8B Instruct's context window?
256,000 tokens, roughly 384 pages of text in a single request.
Does Qwen: Qwen3 VL 8B Instruct support tool calling?
It supports tool calling, structured output.
Can Qwen: Qwen3 VL 8B Instruct process images?
Yes, it accepts image input.