NVIDIA

llama-nemotron-embed-vl-1b-v2

Generally availableOpen weights33K context

Summary

llama-nemotron-embed-vl-1b-v2 is an open-weight language model from NVIDIA, released on 10 Feb 2026. It accepts text and images and generates text, with a 33K-token context window and up to 2K output tokens per response.

Embedding model for semantic search, retrieval, clustering, and ranking pipelines

Last verified 4 Oct 2026. Every figure on this page links to its source.

Specifications

DeveloperNVIDIA
Release date10 Feb 2026 [models.dev]
StatusGenerally available [models.dev]
WeightsOpen weights [models.dev]
Context window33K (32,768 tokens) [models.dev]
Max output2K tokens [models.dev]
InputText and images [models.dev]
OutputText [models.dev]
Reasoning modeNo [models.dev]
Tool callingNo [models.dev]

llama-nemotron-embed-vl-1b-v2 API pricing by provider

USD per million tokens. Sorted by input price.

ProviderInputOutputCached inputContextSource
NVIDIA NIM
nvidia/llama-nemotron-embed-vl-1b-v2
FreeFree–33Kmodels.dev, 4 Oct 2026

Frequently asked questions

What is the context window of llama-nemotron-embed-vl-1b-v2?

llama-nemotron-embed-vl-1b-v2 has a context window of 32,768 tokens (33K) and can generate up to 2,048 tokens in a single response, according to models.dev.

Is llama-nemotron-embed-vl-1b-v2 open source?

llama-nemotron-embed-vl-1b-v2 is an open-weight model. Check the license terms before commercial use.

When was llama-nemotron-embed-vl-1b-v2 released?

NVIDIA released llama-nemotron-embed-vl-1b-v2 on 10 Feb 2026, according to models.dev.

Change history

No changes detected since we started tracking this model. We re-check every source daily.

Sources