meta-llama

Meta: Llama 3.3 70B Instruct

This AI model from Meta is called Llama 3.3 70B Instruct. It processes text input and can handle long contexts up to 131072 tokens. It also supports the use of tools within its interface. For those seeking a balance between cost and performance, this model may be worth considering. Its blended benchmark score of 7.8 across one independent assessment is not exceptional but neither is it low, and pricing starts at $0.1 per million input tokens, with an additional charge for output tokens at $0.32 per million.

Open weights 70.6B parameters, about 39.5 GiB at Q4_K_M: fits comfortably on 0 of 18 consumer graphics cards in the PicksByCard catalog. See which cards run it →
Quality Score
86/100
price + capability + benchmarks
Input Price
$0.10
per 1M tokens · checked 2026-09-23
Output Price
$0.32
per 1M tokens · checked 2026-09-23
Context Window
131,072
tokens

Benchmark results

Independent, published benchmarks. Blended score 7.2 across 1 benchmark, last refreshed 2026-09-23. How scoring works →

BenchmarkMeasuresScore
AI Index broad capability composite 12.6
Model ID
meta-llama/llama-3.3-70b-instruct
Vendor
meta-llama
Released
December 2024
Tokenizer
Llama3
Input Modalities
text
Output Modalities
text
Max Output
16,384 tokens
Tool Calling
✓ supported
Structured Output
✓ supported
Reasoning Mode
not supported
Vision
text only
Audio
no
Moderated
no

What it costs in practice

Computed from the current $0.10/M input and $0.32/M output rates. Run your own numbers →

JobTokensCost
Summarize a 50-page report 30k in / 1.5k out under $0.01
Classify 1,000 customer emails 500k in / 50k out $0.07
A month of a busy support chatbot 5M in / 2M out $1.14

Price & spec history

Tracked daily by PicksByModel since 2026-07-17.

DateInput /MOutput /MContext
2026-09-03 $0.10 $0.32 131,072
2026-08-26 $0.71 $0.71 131,072
2026-08-04 $0.10 $0.32 131,072
2026-07-17 $0.13 $0.40 131,072

Similar models

Quick answers

How much does Meta: Llama 3.3 70B Instruct cost?
$0.10 per million input tokens and $0.32 per million output tokens.
What is Meta: Llama 3.3 70B Instruct's context window?
131,072 tokens, roughly 196 pages of text in a single request.
Does Meta: Llama 3.3 70B Instruct support tool calling?
It supports tool calling, structured output.
Can Meta: Llama 3.3 70B Instruct process images?
No, it is text-only on the input side.

The Model Movers Report

One email every Friday, built from this site's own rankings: the current top five by benchmark score, every model released in the last seven days, and one note worked out from that week's numbers. You can unsubscribe from any issue with one click.