nvidia

NVIDIA: Nemotron Nano 12B 2 VL (free)

Nemotron Nano 12B 2 VL is a multimodal model from NVIDIA that accepts text, images, and video as input, with a context window of 128,000 tokens and a matching 128,000-token completion limit. It supports tool use and reasoning, which makes it applicable to agentic workflows and multi-step tasks. Structured output support is not confirmed. The model is currently free to use with no input or output costs. The free pricing makes it worth shortlisting for developers who want to experiment with video-capable multimodal pipelines without committing a budget. That said, there is no independent benchmark coverage available yet, so its relative performance against comparable models is unproven. Teams making a production decision should treat it as a candidate to test directly rather than one with an established track record, and should weigh the cost advantage against that uncertainty.

Quality Score
89/100
price + capability + benchmarks
Input Price
Free
per 1M tokens
Output Price
Free
per 1M tokens
Context Window
128,000
tokens
Model ID
nvidia/nemotron-nano-12b-v2-vl:free
Vendor
nvidia
Released
October 2025
Tokenizer
Other
Input Modalities
image, text, video
Output Modalities
text
Max Output
128,000 tokens
Tool Calling
✓ supported
Structured Output
not supported
Reasoning Mode
✓ supported
Vision
✓ accepts images
Audio
no
Moderated
no

Similar models

Quick answers

Is NVIDIA: Nemotron Nano 12B 2 VL (free) free to use?
Yes. It is currently listed at no cost per token via OpenRouter, subject to provider rate limits.
What is NVIDIA: Nemotron Nano 12B 2 VL (free)'s context window?
128,000 tokens, roughly 192 pages of text in a single request.
Does NVIDIA: Nemotron Nano 12B 2 VL (free) support tool calling?
It supports tool calling, a reasoning mode.
Can NVIDIA: Nemotron Nano 12B 2 VL (free) process images?
Yes, it accepts image input.