meta-llama

Meta: Llama Guard 4 12B

The Meta: Llama Guard 4 12B is designed to process both text and image inputs within an extensive context length of 1048576 tokens. While it supports multimodal inputs, this model currently does not include reasoning or structured output capabilities; it also lacks support for tools. Given its pricing at $0.18 per million input and output tokens, Llama Guard 4 is suitable for users requiring text-image handling without the need for advanced reasoning or structured data processing. Until independent benchmarks are available to validate its performance, users should consider this model as part of a broader strategy that may include other models with proven capabilities in benchmarking.

Quality Score
89/100
price + capability + benchmarks
Input Price
$0.18
per 1M tokens
Output Price
$0.18
per 1M tokens
Context Window
1,048,576
tokens
Model ID
meta-llama/llama-guard-4-12b
Vendor
meta-llama
Released
April 2025
Tokenizer
Other
Input Modalities
image, text
Output Modalities
text
Max Output
16,384 tokens
Tool Calling
not supported
Structured Output
✓ supported
Reasoning Mode
not supported
Vision
✓ accepts images
Audio
no
Moderated
no

What it costs in practice

Computed from the current $0.18/M input and $0.18/M output rates. Run your own numbers →

JobTokensCost
Summarize a 50-page report 30k in / 1.5k out under $0.01
Classify 1,000 customer emails 500k in / 50k out $0.10
A month of a busy support chatbot 5M in / 2M out $1.26

Price & spec history

Tracked daily by PicksByModel since 2026-07-17.

DateInput /MOutput /MContext
2026-07-22 $0.18 $0.18 1,048,576
2026-07-17 $0.18 $0.18 163,840

Similar models

Quick answers

How much does Meta: Llama Guard 4 12B cost?
$0.18 per million input tokens and $0.18 per million output tokens.
What is Meta: Llama Guard 4 12B's context window?
1,048,576 tokens, roughly 1,572 pages of text in a single request.
Does Meta: Llama Guard 4 12B support tool calling?
It supports structured output.
Can Meta: Llama Guard 4 12B process images?
Yes, it accepts image input.