# llama-nemotron-embed-vl-1b-v2

> llama-nemotron-embed-vl-1b-v2 is an open-weight language model from NVIDIA, released on 10 Feb 2026. It accepts text and images and generates text, with a 33K-token context window and up to 2K output tokens per response.

Source page: https://www.aimodel.directory/models/llama-nemotron-embed-vl-1b-v2
Last verified: 4 Oct 2026

## Specifications

- **Developer:** NVIDIA
- **Release date:** 10 Feb 2026 (source: models.dev, https://github.com/sst/models.dev/blob/dev/providers/nvidia/models/nvidia/llama-nemotron-embed-vl-1b-v2.toml)
- **Status:** Generally available (source: models.dev, https://github.com/sst/models.dev/blob/dev/providers/nvidia/models/nvidia/llama-nemotron-embed-vl-1b-v2.toml)
- **Weights:** Open weights (source: models.dev, https://github.com/sst/models.dev/blob/dev/providers/nvidia/models/nvidia/llama-nemotron-embed-vl-1b-v2.toml)
- **Context window:** 33K (32,768 tokens) (source: models.dev, https://github.com/sst/models.dev/blob/dev/providers/nvidia/models/nvidia/llama-nemotron-embed-vl-1b-v2.toml)
- **Max output:** 2K tokens (source: models.dev, https://github.com/sst/models.dev/blob/dev/providers/nvidia/models/nvidia/llama-nemotron-embed-vl-1b-v2.toml)
- **Input:** Text and images (source: models.dev, https://github.com/sst/models.dev/blob/dev/providers/nvidia/models/nvidia/llama-nemotron-embed-vl-1b-v2.toml)
- **Output:** Text (source: models.dev, https://github.com/sst/models.dev/blob/dev/providers/nvidia/models/nvidia/llama-nemotron-embed-vl-1b-v2.toml)
- **Reasoning mode:** No (source: models.dev, https://github.com/sst/models.dev/blob/dev/providers/nvidia/models/nvidia/llama-nemotron-embed-vl-1b-v2.toml)
- **Tool calling:** No (source: models.dev, https://github.com/sst/models.dev/blob/dev/providers/nvidia/models/nvidia/llama-nemotron-embed-vl-1b-v2.toml)

## API pricing (USD per 1M tokens)

| Provider | Input | Output | Cached input | Context |
|---|---|---|---|---|
| NVIDIA NIM | Free | Free | – | 33K |

## FAQ

### What is the context window of llama-nemotron-embed-vl-1b-v2?

llama-nemotron-embed-vl-1b-v2 has a context window of 32,768 tokens (33K) and can generate up to 2,048 tokens in a single response, according to models.dev.

### Is llama-nemotron-embed-vl-1b-v2 open source?

llama-nemotron-embed-vl-1b-v2 is an open-weight model. Check the license terms before commercial use.

### When was llama-nemotron-embed-vl-1b-v2 released?

NVIDIA released llama-nemotron-embed-vl-1b-v2 on 10 Feb 2026, according to models.dev.

---
Data licensed CC BY 4.0. Cite as: AI Model Directory, "llama-nemotron-embed-vl-1b-v2", https://www.aimodel.directory/models/llama-nemotron-embed-vl-1b-v2