value rankings

Intelligence Per Dollar

Benchmark points per blended dollar across every credibly ranked model. Scores are min-max normalized across 179 models with 4+ independent benchmarks; cost assumes a typical 3:1 input:output token mix. Leader = 100.

This is not a quality ranking. #1 here means the most benchmark points per dollar, so a mid-scoring budget model will outrank a frontier model that costs 100x more. Check the Score and % of top score columns for raw capability, and use the task rankings when output quality is what compounds in your workflow.

Value leaderboard

#ModelScore% of top scoreIn / Out per 1MBlended $/1MValue
1inclusionAI: Ling-2.6-flash BEST VALUEinclusionai/ling-2.6-flash21.522%$0.01 / $0.03$0.01100.0
2OpenAI: gpt-oss-120bopenai/gpt-oss-120b43.143%$0.04 / $0.17$0.0742.8
3Mistral: Mistral Small 3mistralai/mistral-small-24b-instruct-250128.829%$0.05 / $0.08$0.0634.9
4OpenAI: gpt-oss-20bopenai/gpt-oss-20b25.626%$0.03 / $0.14$0.0631.1
5DeepSeek: DeepSeek V4 Flashdeepseek/deepseek-v4-flash74.574%$0.14 / $0.28$0.1829.7
6Mistral: Mistral Small 3.2 24Bmistralai/mistral-small-3.2-24b-instruct59.660%$0.10 / $0.30$0.1527.7
7Z.ai: GLM 4.7 Flashz-ai/glm-4.7-flash46.346%$0.06 / $0.40$0.1522.3
8Google: Gemma 4 26B A4B google/gemma-4-26b-a4b-it51.852%$0.12 / $0.35$0.1820.4
9StepFun: Step 3.5 Flashstepfun/step-3.5-flash41.942%$0.10 / $0.30$0.1519.5
10Google: Gemma 4 31Bgoogle/gemma-4-31b-it56.957%$0.14 / $0.40$0.2119.4
11Qwen: Qwen3.5-9Bqwen/qwen3.5-9b31.031%$0.10 / $0.15$0.1119.2
12Qwen: Qwen3 32Bqwen/qwen3-32b35.035%$0.08 / $0.28$0.1318.8
13NVIDIA: Nemotron 3 Supernvidia/nemotron-3-super-120b-a12b39.139%$0.09 / $0.40$0.1616.7
14inclusionAI: Ring-2.6-1Tinclusionai/ring-2.6-1t50.751%$0.07 / $0.62$0.2116.6
15Google: Gemma 3 4Bgoogle/gemma-3-4b-it14.514%$0.05 / $0.10$0.0616.2
16OpenAI: GPT-5 Nanoopenai/gpt-5-nano31.632%$0.05 / $0.40$0.1416.0
17DeepSeek: DeepSeek V3deepseek/deepseek-chat78.378%$0.20 / $0.80$0.3515.6
18inclusionAI: Ling-2.6-1Tinclusionai/ling-2.6-1t42.042%$0.07 / $0.62$0.2113.8
19Qwen: Qwen3 Coder 30B A3B Instructqwen/qwen3-coder-30b-a3b-instruct21.021%$0.07 / $0.27$0.1212.2
20Nex AGI: Nex-N2-Pronex-agi/nex-n2-pro72.873%$0.25 / $1.00$0.4411.6
21IBM: Granite 4.1 8Bibm-granite/granite-4.1-8b10.410%$0.05 / $0.10$0.0611.6
22MiniMax: MiniMax M2.5minimax/minimax-m2.554.755%$0.15 / $0.90$0.3411.3
23DeepSeek: DeepSeek V4 Prodeepseek/deepseek-v4-pro81.782%$0.43 / $0.87$0.5410.5
24MiniMax: MiniMax M2.7minimax/minimax-m2.764.564%$0.25 / $1.00$0.4410.3
25Qwen: Qwen3.6 35B A3Bqwen/qwen3.6-35b-a3b52.352%$0.14 / $1.00$0.3510.3
26MiniMax: MiniMax M3minimax/minimax-m377.277%$0.30 / $1.20$0.5210.3
27OpenAI: GPT-5.4 Nanoopenai/gpt-5.4-nano67.568%$0.20 / $1.25$0.4610.2
28Xiaomi: MiMo-V2.5-Proxiaomi/mimo-v2.5-pro73.173%$0.43 / $0.87$0.549.4
29Qwen: Qwen3.5-35B-A3Bqwen/qwen3.5-35b-a3b47.447%$0.14 / $1.00$0.359.3
30Qwen: Qwen3 Coder Nextqwen/qwen3-coder-next35.536%$0.11 / $0.80$0.288.8

Budget champions : 80+ score, cheapest first

#ModelScore% of top scoreIn / Out per 1MBlended $/1MValue
1DeepSeek: DeepSeek V4 Prodeepseek/deepseek-v4-pro81.782%$0.43 / $0.87$0.5410.5
2Z.ai: GLM 5.2z-ai/glm-5.288.789%$0.76 / $2.39$1.175.3
3OpenAI: o4 Mini Highopenai/o4-mini-high81.081%$1.10 / $4.40$1.932.9
4Meta: Muse Spark 1.1meta/muse-spark-1.183.684%$1.25 / $4.25$2.002.9
5OpenAI: GPT-5.6 Lunaopenai/gpt-5.6-luna88.488%$1.00 / $6.00$2.252.7
6Google: Gemini 3.6 Flashgoogle/gemini-3.6-flash83.984%$1.50 / $7.50$3.002.0
7xAI: Grok 4.5x-ai/grok-4.590.290%$2.00 / $6.00$3.002.1
8Google: Gemini 3.5 Flashgoogle/gemini-3.5-flash83.383%$1.50 / $9.00$3.381.7
9Google: Gemini 2.5 Progoogle/gemini-2.5-pro94.294%$1.25 / $10.00$3.441.9
10Anthropic: Claude Sonnet 5anthropic/claude-sonnet-590.490%$2.00 / $10.00$4.001.6

Free models with credible scores

Per-dollar math breaks at $0. These are simply the strongest free options:

Assumptions

Value = blended benchmark score divided by blended price per million tokens, indexed to the leader. A 3:1 input:output ratio fits most chat and RAG workloads; estimate your exact mix with the cost calculator. Scoring details in the methodology. Models with fewer than 4 independent benchmarks are excluded rather than guessed.