google

Google: Gemma 4 31B

The Google Gemma 4 31B model is a versatile AI system that can process and generate content from multiple modalities, including text, images, and video. It has a large context length of 262144 tokens, allowing it to handle complex and nuanced inputs. For users who need a reliable AI solution with advanced reasoning capabilities, the Google Gemma 4 31B may be worth considering. Its pricing is competitive at $0.09 per million input tokens and $0.34 per million output tokens. The model's blended benchmark score of 56.5 across four independent benchmarks suggests it performs well in various tasks. However, its high price point may deter some users who require a more affordable solution. Those with a budget to spare or specific needs that align with this model's strengths may find the Google Gemma 4 31B an attractive option.

Open weights 31.3B parameters, about 17.5 GiB at Q4_K_M: fits comfortably on 4 of 17 consumer graphics cards in the PicksByCard catalog. See which cards run it →
Quality Score
100/100
price + capability + benchmarks
Input Price
$0.09
per 1M tokens · checked 2026-09-04
Output Price
$0.34
per 1M tokens · checked 2026-09-04
Context Window
262,144
tokens

Benchmark results

Independent, published benchmarks. Blended score 56.5 across 4 benchmarks, last refreshed 2026-09-04. How scoring works →

BenchmarkMeasuresScore
AI Index broad capability composite 49.0
AI Index Coding software engineering tasks 71.7
AI Index Agentic multi-step tool-using tasks 23.7
EQ-Bench emotional understanding in dialogue 70.8
Model ID
google/gemma-4-31b-it
Vendor
google
Released
April 2026
Tokenizer
Gemma
Input Modalities
image, text, video
Output Modalities
text
Max Output
16,384 tokens
Tool Calling
✓ supported
Structured Output
✓ supported
Reasoning Mode
✓ supported
Vision
✓ accepts images
Audio
no
Moderated
no

What it costs in practice

Computed from the current $0.09/M input and $0.34/M output rates. Run your own numbers →

JobTokensCost
Summarize a 50-page report 30k in / 1.5k out under $0.01
Classify 1,000 customer emails 500k in / 50k out $0.06
A month of a busy support chatbot 5M in / 2M out $1.13

Price & spec history

Tracked daily by PicksByModel since 2026-07-17.

DateInput /MOutput /MContext
2026-08-26 $0.09 $0.34 262,144
2026-08-21 $0.10 $0.34 262,144
2026-08-19 $0.09 $0.34 262,144
2026-07-30 $0.10 $0.34 262,144
2026-07-25 $0.14 $0.40 262,144
2026-07-24 $0.10 $0.35 262,144
2026-07-23 $0.12 $0.35 262,144
2026-07-21 $0.12 $0.37 262,144
2026-07-17 $0.22 $0.55 262,144

Category rankings

Where Google: Gemma 4 31B places across the 5 categories it ranks in. How we rank →

#CategoryScore
#13 Real-Time ChatLatency · of 25 ranked 118
#14 Self-Hosted / LocalCost · of 25 ranked 117
#18 Cheap Bulk InferenceCost · of 25 ranked 137
#22 Social Media PostsWriting · of 25 ranked 119
#22 Voice Assistant BackendVoice · of 25 ranked 123

Similar models

Quick answers

How much does Google: Gemma 4 31B cost?
$0.09 per million input tokens and $0.34 per million output tokens.
What is Google: Gemma 4 31B's context window?
262,144 tokens, roughly 393 pages of text in a single request.
Does Google: Gemma 4 31B support tool calling?
It supports tool calling, structured output, a reasoning mode.
Can Google: Gemma 4 31B process images?
Yes, it accepts image input.