value rankings

Intelligence Per Dollar

Benchmark points per blended dollar across every credibly ranked model. Scores are min-max normalized across 254 models with 4+ independent benchmarks; cost assumes a typical 3:1 input:output token mix. Leader = 100.

This is not a quality ranking. #1 here means the most benchmark points per dollar, so a mid-scoring budget model will outrank a frontier model that costs 100x more. Check the Score and % of top score columns for raw capability, and use the task rankings when output quality is what compounds in your workflow.

Value leaderboard

#ModelScore% of top scoreIn / Out per 1MBlended $/1MValue
1inclusionAI: Ling 3.0 Flash BEST VALUEinclusionai/ling-3.0-flash39.339%$0.02 / $0.06$0.03100.0
2Z.ai: GLM 5.3 Flash (batch)z-ai/glm-5.3-flash:batch70.671%$0.06 / $0.20$0.1059.6
3OpenAI: gpt-oss-20bopenai/gpt-oss-20b24.324%$0.02 / $0.09$0.0454.1
4OpenAI: GPT-6 Luna (batch)openai/gpt-6-luna:batch62.262%$0.05 / $0.25$0.1049.9
5OpenAI: gpt-oss-20b (batch)openai/gpt-oss-20b:batch24.324%$0.02 / $0.11$0.0542.3
6Mistral: Mistral Small 3mistralai/mistral-small-24b-instruct-250128.829%$0.05 / $0.08$0.0640.1
7Mistral: Mistral Small 3.2 24Bmistralai/mistral-small-3.2-24b-instruct59.660%$0.09 / $0.25$0.1336.0
8inclusionAI: Ling 3.0 Flash VLinclusionai/ling-3.0-flash-vl38.639%$0.06 / $0.18$0.0934.4
9DeepSeek: DeepSeek V4.1 Flash (batch)deepseek/deepseek-v4.1-flash:batch66.366%$0.11 / $0.34$0.1731.6
10inclusionAI: Ling 3.0 Flash Fininclusionai/ling-3.0-flash-fin35.035%$0.06 / $0.18$0.0931.2
11Google: Gemma 4 26B A4B google/gemma-4-26b-a4b-it50.951%$0.09 / $0.30$0.1428.6
12Google: Gemma 4 31Bgoogle/gemma-4-31b-it53.754%$0.09 / $0.34$0.1528.2
13Google: Gemma 3 4Bgoogle/gemma-3-4b-it21.321%$0.05 / $0.10$0.0627.3
14OpenAI: GPT-6 Lunaopenai/gpt-6-luna62.262%$0.10 / $0.50$0.2024.9
15Z.ai: GLM 5.3 Flashz-ai/glm-5.3-flash70.671%$0.15 / $0.50$0.2423.8
16Upstage: Solar Pro 4upstage/solar-pro445.345%$0.09 / $0.36$0.1623.1
17OpenAI: GPT-5.6 Luna (batch)openai/gpt-5.6-luna:batch62.362%$0.10 / $0.60$0.2322.2
18Z.ai: GLM 4.7 Flashz-ai/glm-4.7-flash37.838%$0.06 / $0.40$0.1520.8
19Qwen: Qwen3 32Bqwen/qwen3-32b33.133%$0.08 / $0.28$0.1320.4
20DeepSeek: DeepSeek V4.1 Flashdeepseek/deepseek-v4.1-flash66.366%$0.15 / $0.60$0.2620.2
21OpenAI: GPT-5 Nano (batch)openai/gpt-5-nano:batch17.117%$0.03 / $0.20$0.0719.9
22OpenAI: GPT-4.1 Nano (batch)openai/gpt-4.1-nano:batch16.717%$0.05 / $0.20$0.0915.3
23DeepSeek: DeepSeek V4 Flash Vision Expdeepseek/deepseek-v4-flash-vision-exp57.758%$0.22 / $0.66$0.3314.0
24NVIDIA: Nemotron 3.5 Lightningnvidia/nemotron-3.5-lightning16.917%$0.07 / $0.20$0.1013.2
25StepFun: Step 3.5 Flashstepfun/step-3.5-flash24.524%$0.10 / $0.30$0.1513.1
26Qwen: Qwen3.5-9Bqwen/qwen3.5-9b18.318%$0.10 / $0.15$0.1113.0
27DeepSeek: DeepSeek V3deepseek/deepseek-chat70.470%$0.32 / $0.89$0.4612.2
28OpenAI: gpt-oss-120bopenai/gpt-oss-120b39.239%$0.15 / $0.60$0.2612.0
29Xiaomi: MiMo-V2.6-Proxiaomi/mimo-v2.6-pro79.079%$0.43 / $0.87$0.5411.6
30OpenAI: GPT-5.6 Lunaopenai/gpt-5.6-luna62.362%$0.20 / $1.20$0.4511.1

Budget champions : 80+ score, cheapest first

#ModelScore% of top scoreIn / Out per 1MBlended $/1MValue
1Google: Gemini 2.5 Pro (batch)google/gemini-2.5-pro:batch94.294%$0.62 / $5.00$1.724.4
2OpenAI: o4 Mini Highopenai/o4-mini-high81.081%$1.10 / $4.40$1.933.4
3OpenAI: GPT-6 Sol (batch)openai/gpt-6-sol:batch81.281%$1.00 / $5.00$2.003.3
4Meta: Muse Spark 1.3meta/muse-spark-1.382.382%$1.25 / $4.25$2.003.3
5OpenAI: GPT-5.6 Sol (batch)openai/gpt-5.6-sol:batch80.280%$1.00 / $5.00$2.003.2
6OpenAI: GPT-5.4 (batch)openai/gpt-5.4:batch80.981%$1.25 / $7.50$2.812.3
7Google: Gemini 2.5 Progoogle/gemini-2.5-pro94.294%$1.25 / $10.00$3.442.2
8OpenAI: GPT-6 Solopenai/gpt-6-sol81.281%$2.00 / $10.00$4.001.6
9Anthropic: Claude Opus 5.5 (batch)anthropic/claude-opus-5.5:batch100.0100%$2.00 / $10.00$4.002.0
10OpenAI: GPT-5.6 Solopenai/gpt-5.6-sol80.280%$2.00 / $10.00$4.001.6

Free models with credible scores

Per-dollar math breaks at $0. These are simply the strongest free options:

Assumptions

Value = blended benchmark score divided by blended price per million tokens, indexed to the leader. A 3:1 input:output ratio fits most chat and RAG workloads; estimate your exact mix with the cost calculator. Scoring details in the methodology. Models with fewer than 4 independent benchmarks are excluded rather than guessed.

The Model Movers Report

One email every Friday, built from this site's own rankings: the current top five by benchmark score, every model released in the last seven days, and one note worked out from that week's numbers. You can unsubscribe from any issue with one click.