inclusionai

inclusionAI: Ling 3.0 Flash VL

The inclusionAI: Ling 3.0 Flash VL is a multimodal model that can process text, images, and video inputs simultaneously. It supports context lengths of up to 131072 tokens and has a maximum completion token limit of 32768. Additionally, it comes with built-in tools and reasoning capabilities. For users who need to handle multimedia inputs and require advanced reasoning, this model may be worth shortlisting due to its robust feature set. However, its pricing is relatively high at $0.06 per million input tokens and $0.18 per million output tokens. With a blended benchmark score of 38.6 across one independent benchmark, it sits in the middle ground among models with similar capabilities; further evaluation may be necessary to determine its effectiveness in specific use cases.

Quality Score
100/100
price + capability + benchmarks
Input Price
$0.06
per 1M tokens · checked 2026-09-23
Output Price
$0.18
per 1M tokens · checked 2026-09-23
Context Window
131,072
tokens

Benchmark results

Independent, published benchmarks. Blended score 38.6 across 1 benchmark, last refreshed 2026-09-23. How scoring works →

BenchmarkMeasuresScore
AI Index broad capability composite 40.5
Model ID
inclusionai/ling-3.0-flash-vl
Vendor
inclusionai
Released
September 2026
Tokenizer
Other
Input Modalities
text, image, video
Output Modalities
text
Max Output
32,768 tokens
Tool Calling
✓ supported
Structured Output
✓ supported
Reasoning Mode
✓ supported
Vision
✓ accepts images
Audio
no
Moderated
no

What it costs in practice

Computed from the current $0.06/M input and $0.18/M output rates. Run your own numbers →

JobTokensCost
Summarize a 50-page report 30k in / 1.5k out under $0.01
Classify 1,000 customer emails 500k in / 50k out $0.04
A month of a busy support chatbot 5M in / 2M out $0.66

Strong choice for

Category rankings

Where inclusionAI: Ling 3.0 Flash VL places across the 10 categories it ranks in. How we rank →

#CategoryScore
#1 Social Media PostsWriting · of 25 ranked 120
#1 Voice Assistant BackendVoice · of 25 ranked 124
#4 Real-Time ChatLatency · of 25 ranked 118
#5 Video Auto-TaggingVideo · of 25 ranked 123
#6 Cheap Bulk InferenceCost · of 25 ranked 137
#10 Self-Hosted / LocalCost · of 25 ranked 117
#11 Customer SupportBusiness · of 25 ranked 128
#12 JSON ExtractionData · of 25 ranked 132
#19 Bulk Data LabelingData · of 25 ranked 129
#20 Dataset AnnotationResearch · of 25 ranked 133

Similar models

Quick answers

How much does inclusionAI: Ling 3.0 Flash VL cost?
$0.06 per million input tokens and $0.18 per million output tokens.
What is inclusionAI: Ling 3.0 Flash VL's context window?
131,072 tokens, roughly 196 pages of text in a single request.
Does inclusionAI: Ling 3.0 Flash VL support tool calling?
It supports tool calling, structured output, a reasoning mode.
Can inclusionAI: Ling 3.0 Flash VL process images?
Yes, it accepts image input.

The Model Movers Report

One email every Friday, built from this site's own rankings: the current top five by benchmark score, every model released in the last seven days, and one note worked out from that week's numbers. You can unsubscribe from any issue with one click.