Vision · best for

Top picks for Image Captioning (2026)

Accessible alt text and detailed image descriptions. Ranked from 422 live models on the OpenRouter catalog, weighted for vision input, low latency.

Updated 2026-09-04 · prices checked at this morning's rebuild

What this is Ranked by capability match + real benchmark scores (Aider Polyglot, Artificial Analysis Intelligence Index) + live pricing. Models need the right specs for Image Captioning, then benchmark performance refines the order. Full methodology →

Which should you use? Google: Gemini 3.8 Flash tops this ranking on blended score. If cost drives the decision, Z.ai: GLM 5.3 Flash is the cheapest of the leaders at $0.07/M input.

#ModelScoreIn / 1MOut / 1MContext
1 Google: Gemini 3.8 Flashgoogle/gemini-3.8-flash 121 $0.75 $3.75 1,048,576 Details →
2 Google: Gemini 3.8 Flash (batch)google/gemini-3.8-flash:batch 121 $0.38 $1.88 1,048,576 Details →
3 Z.ai: GLM 5.3 Flashz-ai/glm-5.3-flash 121 $0.07 $0.25 1,310,720 Details →
4 Z.ai: GLM 5.3 Flash (batch)z-ai/glm-5.3-flash:batch 121 $0.15 $0.50 1,048,575 Details →
5 Google: Gemini 3.7 Flashgoogle/gemini-3.7-flash 121 $0.75 $3.75 1,048,576 Details →
6 Google: Gemini 3.7 Flash (batch)google/gemini-3.7-flash:batch 121 $0.38 $1.88 1,048,576 Details →
7 Qwen: Qwen3.8 27Bqwen/qwen3.8-27b 121 $0.42 $3.00 1,000,000 Details →
8 Google: Gemini 3.6 Flashgoogle/gemini-3.6-flash 121 $0.75 $3.75 1,048,576 Details →
9 Google: Gemini 3.6 Flash (batch)google/gemini-3.6-flash:batch 121 $0.38 $1.88 1,048,576 Details →
10 OpenAI: GPT-5.6 Lunaopenai/gpt-5.6-luna 121 $0.20 $1.20 1,050,000 Details →
11 OpenAI: GPT-5.6 Luna (batch)openai/gpt-5.6-luna:batch 121 $0.10 $0.60 1,050,000 Details →
12 Google: Gemini 3.5 Flash (batch)google/gemini-3.5-flash:batch 121 $0.75 $4.50 1,048,576 Details →
13 DeepSeek: DeepSeek V4 Flash Vision Expdeepseek/deepseek-v4-flash-vision-exp 121 $0.22 $0.66 1,048,576 Details →
14 MiniMax: MiniMax M3minimax/minimax-m3 121 $0.30 $1.20 1,048,576 Details →
15 MiniMax: MiniMax M3 (batch)minimax/minimax-m3:batch 121 $0.30 $1.20 524,288 Details →
AI Photo Editing Retouch4me AI retouching that finishes what the generator started: skin, eyes, dust, and color in one pass.
Try free →

Affiliate link. PicksByModel may earn a commission at no extra cost to you.

How we ranked these

For Image Captioning, we weight models on vision input, low latency. Scores combine each model's public specs with independent benchmark results (Aider Polyglot coding scores, Artificial Analysis intelligence/coding/agentic indices) and live pricing. See full methodology →

About Image Captioning

Image captioning is the task of generating natural language descriptions for images, producing text that conveys visual content accurately and contextually. Use this when you need accessible alt text for web content, searchable descriptions for image archives, or automated tagging for large visual datasets. Good models balance accuracy with brevity, describing objects and relationships without hallucinating details that aren't present, while poor ones produce generic or misleading text. The critical trade-off: vision-language models like BLIP or LLaVA generate more natural captions than older CNN-based approaches but require significantly more computational resources, typically 2-4x slower inference time depending on model size.

When to use: Use this when you need to automatically generate text descriptions for images so they're readable by screen readers, searchable in databases, or accessible to people who can't see them.

Common questions

Which AI model produces the most accurate image captions today?

BLIP-2 and LLaVA represent the current best-in-class for caption quality, with LLaVA-1.6 offering particularly strong reasoning about image relationships. If you need faster inference, BLIP (the original) still delivers solid accuracy at half the computational cost. For production use, your choice depends on whether you prioritize caption quality or response latency.

How much does it cost to caption thousands of images at scale?

Running open-source models like LLaVA yourself costs roughly $0.0001-0.0005 per image on cloud compute, while API services like Google Vision or AWS Rekognition charge $0.0015-0.004 per image. For 10,000 images, self-hosted models save 50-70% but require infrastructure setup, whereas APIs eliminate operational overhead.

Related tasks