llama-nemotron-embed-vl-1b-v2
Summary
llama-nemotron-embed-vl-1b-v2 is an open-weight language model from NVIDIA, released on 10 Feb 2026. It accepts text and images and generates text, with a 33K-token context window and up to 2K output tokens per response.
Embedding model for semantic search, retrieval, clustering, and ranking pipelines
Last verified 4 Oct 2026. Every figure on this page links to its source.
Specifications
| Developer | NVIDIA |
|---|---|
| Release date | 10 Feb 2026 [models.dev] |
| Status | Generally available [models.dev] |
| Weights | Open weights [models.dev] |
| Context window | 33K (32,768 tokens) [models.dev] |
| Max output | 2K tokens [models.dev] |
| Input | Text and images [models.dev] |
| Output | Text [models.dev] |
| Reasoning mode | No [models.dev] |
| Tool calling | No [models.dev] |
llama-nemotron-embed-vl-1b-v2 API pricing by provider
USD per million tokens. Sorted by input price.
| Provider | Input | Output | Cached input | Context | Source |
|---|---|---|---|---|---|
| NVIDIA NIM nvidia/llama-nemotron-embed-vl-1b-v2 | Free | Free | – | 33K | models.dev, 4 Oct 2026 |
Frequently asked questions
What is the context window of llama-nemotron-embed-vl-1b-v2?
llama-nemotron-embed-vl-1b-v2 has a context window of 32,768 tokens (33K) and can generate up to 2,048 tokens in a single response, according to models.dev.
Is llama-nemotron-embed-vl-1b-v2 open source?
llama-nemotron-embed-vl-1b-v2 is an open-weight model. Check the license terms before commercial use.
When was llama-nemotron-embed-vl-1b-v2 released?
NVIDIA released llama-nemotron-embed-vl-1b-v2 on 10 Feb 2026, according to models.dev.
Change history
No changes detected since we started tracking this model. We re-check every source daily.