Intelligence Per Dollar
Benchmark points per blended dollar across every credibly ranked model. Scores are min-max normalized across 179 models with 4+ independent benchmarks; cost assumes a typical 3:1 input:output token mix. Leader = 100.
This is not a quality ranking. #1 here means the most benchmark points per dollar, so a mid-scoring budget model will outrank a frontier model that costs 100x more. Check the Score and % of top score columns for raw capability, and use the task rankings when output quality is what compounds in your workflow.
Value leaderboard
| # | Model | Score | % of top score | In / Out per 1M | Blended $/1M | Value |
|---|---|---|---|---|---|---|
| 1 | inclusionAI: Ling-2.6-flash BEST VALUEinclusionai/ling-2.6-flash | 21.5 | 22% | $0.01 / $0.03 | $0.01 | 100.0 |
| 2 | OpenAI: gpt-oss-120bopenai/gpt-oss-120b | 43.1 | 43% | $0.04 / $0.17 | $0.07 | 42.8 |
| 3 | Mistral: Mistral Small 3mistralai/mistral-small-24b-instruct-2501 | 28.8 | 29% | $0.05 / $0.08 | $0.06 | 34.9 |
| 4 | OpenAI: gpt-oss-20bopenai/gpt-oss-20b | 25.6 | 26% | $0.03 / $0.14 | $0.06 | 31.1 |
| 5 | DeepSeek: DeepSeek V4 Flashdeepseek/deepseek-v4-flash | 74.5 | 74% | $0.14 / $0.28 | $0.18 | 29.7 |
| 6 | Mistral: Mistral Small 3.2 24Bmistralai/mistral-small-3.2-24b-instruct | 59.6 | 60% | $0.10 / $0.30 | $0.15 | 27.7 |
| 7 | Z.ai: GLM 4.7 Flashz-ai/glm-4.7-flash | 46.3 | 46% | $0.06 / $0.40 | $0.15 | 22.3 |
| 8 | Google: Gemma 4 26B A4B google/gemma-4-26b-a4b-it | 51.8 | 52% | $0.12 / $0.35 | $0.18 | 20.4 |
| 9 | StepFun: Step 3.5 Flashstepfun/step-3.5-flash | 41.9 | 42% | $0.10 / $0.30 | $0.15 | 19.5 |
| 10 | Google: Gemma 4 31Bgoogle/gemma-4-31b-it | 56.9 | 57% | $0.14 / $0.40 | $0.21 | 19.4 |
| 11 | Qwen: Qwen3.5-9Bqwen/qwen3.5-9b | 31.0 | 31% | $0.10 / $0.15 | $0.11 | 19.2 |
| 12 | Qwen: Qwen3 32Bqwen/qwen3-32b | 35.0 | 35% | $0.08 / $0.28 | $0.13 | 18.8 |
| 13 | NVIDIA: Nemotron 3 Supernvidia/nemotron-3-super-120b-a12b | 39.1 | 39% | $0.09 / $0.40 | $0.16 | 16.7 |
| 14 | inclusionAI: Ring-2.6-1Tinclusionai/ring-2.6-1t | 50.7 | 51% | $0.07 / $0.62 | $0.21 | 16.6 |
| 15 | Google: Gemma 3 4Bgoogle/gemma-3-4b-it | 14.5 | 14% | $0.05 / $0.10 | $0.06 | 16.2 |
| 16 | OpenAI: GPT-5 Nanoopenai/gpt-5-nano | 31.6 | 32% | $0.05 / $0.40 | $0.14 | 16.0 |
| 17 | DeepSeek: DeepSeek V3deepseek/deepseek-chat | 78.3 | 78% | $0.20 / $0.80 | $0.35 | 15.6 |
| 18 | inclusionAI: Ling-2.6-1Tinclusionai/ling-2.6-1t | 42.0 | 42% | $0.07 / $0.62 | $0.21 | 13.8 |
| 19 | Qwen: Qwen3 Coder 30B A3B Instructqwen/qwen3-coder-30b-a3b-instruct | 21.0 | 21% | $0.07 / $0.27 | $0.12 | 12.2 |
| 20 | Nex AGI: Nex-N2-Pronex-agi/nex-n2-pro | 72.8 | 73% | $0.25 / $1.00 | $0.44 | 11.6 |
| 21 | IBM: Granite 4.1 8Bibm-granite/granite-4.1-8b | 10.4 | 10% | $0.05 / $0.10 | $0.06 | 11.6 |
| 22 | MiniMax: MiniMax M2.5minimax/minimax-m2.5 | 54.7 | 55% | $0.15 / $0.90 | $0.34 | 11.3 |
| 23 | DeepSeek: DeepSeek V4 Prodeepseek/deepseek-v4-pro | 81.7 | 82% | $0.43 / $0.87 | $0.54 | 10.5 |
| 24 | MiniMax: MiniMax M2.7minimax/minimax-m2.7 | 64.5 | 64% | $0.25 / $1.00 | $0.44 | 10.3 |
| 25 | Qwen: Qwen3.6 35B A3Bqwen/qwen3.6-35b-a3b | 52.3 | 52% | $0.14 / $1.00 | $0.35 | 10.3 |
| 26 | MiniMax: MiniMax M3minimax/minimax-m3 | 77.2 | 77% | $0.30 / $1.20 | $0.52 | 10.3 |
| 27 | OpenAI: GPT-5.4 Nanoopenai/gpt-5.4-nano | 67.5 | 68% | $0.20 / $1.25 | $0.46 | 10.2 |
| 28 | Xiaomi: MiMo-V2.5-Proxiaomi/mimo-v2.5-pro | 73.1 | 73% | $0.43 / $0.87 | $0.54 | 9.4 |
| 29 | Qwen: Qwen3.5-35B-A3Bqwen/qwen3.5-35b-a3b | 47.4 | 47% | $0.14 / $1.00 | $0.35 | 9.3 |
| 30 | Qwen: Qwen3 Coder Nextqwen/qwen3-coder-next | 35.5 | 36% | $0.11 / $0.80 | $0.28 | 8.8 |
Budget champions : 80+ score, cheapest first
| # | Model | Score | % of top score | In / Out per 1M | Blended $/1M | Value |
|---|---|---|---|---|---|---|
| 1 | DeepSeek: DeepSeek V4 Prodeepseek/deepseek-v4-pro | 81.7 | 82% | $0.43 / $0.87 | $0.54 | 10.5 |
| 2 | Z.ai: GLM 5.2z-ai/glm-5.2 | 88.7 | 89% | $0.76 / $2.39 | $1.17 | 5.3 |
| 3 | OpenAI: o4 Mini Highopenai/o4-mini-high | 81.0 | 81% | $1.10 / $4.40 | $1.93 | 2.9 |
| 4 | Meta: Muse Spark 1.1meta/muse-spark-1.1 | 83.6 | 84% | $1.25 / $4.25 | $2.00 | 2.9 |
| 5 | OpenAI: GPT-5.6 Lunaopenai/gpt-5.6-luna | 88.4 | 88% | $1.00 / $6.00 | $2.25 | 2.7 |
| 6 | Google: Gemini 3.6 Flashgoogle/gemini-3.6-flash | 83.9 | 84% | $1.50 / $7.50 | $3.00 | 2.0 |
| 7 | xAI: Grok 4.5x-ai/grok-4.5 | 90.2 | 90% | $2.00 / $6.00 | $3.00 | 2.1 |
| 8 | Google: Gemini 3.5 Flashgoogle/gemini-3.5-flash | 83.3 | 83% | $1.50 / $9.00 | $3.38 | 1.7 |
| 9 | Google: Gemini 2.5 Progoogle/gemini-2.5-pro | 94.2 | 94% | $1.25 / $10.00 | $3.44 | 1.9 |
| 10 | Anthropic: Claude Sonnet 5anthropic/claude-sonnet-5 | 90.4 | 90% | $2.00 / $10.00 | $4.00 | 1.6 |
Free models with credible scores
Per-dollar math breaks at $0. These are simply the strongest free options:
- NVIDIA: Nemotron 3 Ultra (free) : score 63.5
- Google: Gemma 4 31B (free) : score 56.9
- Google: Gemma 4 26B A4B (free) : score 51.8
- NVIDIA: Nemotron 3 Super (free) : score 39.1
- OpenAI: gpt-oss-20b (free) : score 25.6
- NVIDIA: Nemotron 3 Nano 30B A3B (free) : score 10.6
Assumptions
Value = blended benchmark score divided by blended price per million tokens, indexed to the leader. A 3:1 input:output ratio fits most chat and RAG workloads; estimate your exact mix with the cost calculator. Scoring details in the methodology. Models with fewer than 4 independent benchmarks are excluded rather than guessed.