# Llama-4-Maverick-17B-128E-Instruct-FP8

> Llama-4-Maverick-17B-128E-Instruct-FP8 is an open-weight language model from Meta, released on 5 Apr 2025. It accepts text and images and generates text, with a 128K-token context window and up to 4K output tokens per response.

Source page: https://www.aimodel.directory/models/llama-4-maverick-17b-128e-instruct-fp8
Last verified: 4 Oct 2026

## Specifications

- **Developer:** Meta
- **Release date:** 5 Apr 2025 (source: models.dev, https://github.com/sst/models.dev/blob/dev/providers/llama/models/llama-4-maverick-17b-128e-instruct-fp8.toml)
- **Status:** Generally available (source: models.dev, https://github.com/sst/models.dev/blob/dev/providers/llama/models/llama-4-maverick-17b-128e-instruct-fp8.toml)
- **Weights:** Open weights (source: models.dev, https://github.com/sst/models.dev/blob/dev/providers/llama/models/llama-4-maverick-17b-128e-instruct-fp8.toml)
- **Context window:** 128K (128,000 tokens) (source: models.dev, https://github.com/sst/models.dev/blob/dev/providers/llama/models/llama-4-maverick-17b-128e-instruct-fp8.toml)
- **Max output:** 4K tokens (source: models.dev, https://github.com/sst/models.dev/blob/dev/providers/llama/models/llama-4-maverick-17b-128e-instruct-fp8.toml)
- **Knowledge cutoff:** Aug 2024 (source: models.dev, https://github.com/sst/models.dev/blob/dev/providers/llama/models/llama-4-maverick-17b-128e-instruct-fp8.toml)
- **Input:** Text and images (source: models.dev, https://github.com/sst/models.dev/blob/dev/providers/llama/models/llama-4-maverick-17b-128e-instruct-fp8.toml)
- **Output:** Text (source: models.dev, https://github.com/sst/models.dev/blob/dev/providers/llama/models/llama-4-maverick-17b-128e-instruct-fp8.toml)
- **Reasoning mode:** No (source: models.dev, https://github.com/sst/models.dev/blob/dev/providers/llama/models/llama-4-maverick-17b-128e-instruct-fp8.toml)
- **Tool calling:** Yes (source: models.dev, https://github.com/sst/models.dev/blob/dev/providers/llama/models/llama-4-maverick-17b-128e-instruct-fp8.toml)

## API pricing (USD per 1M tokens)

| Provider | Input | Output | Cached input | Context |
|---|---|---|---|---|
| Llama API | Free | Free | – | 128K |

## FAQ

### What is the context window of Llama-4-Maverick-17B-128E-Instruct-FP8?

Llama-4-Maverick-17B-128E-Instruct-FP8 has a context window of 128,000 tokens (128K) and can generate up to 4,096 tokens in a single response, according to models.dev.

### Is Llama-4-Maverick-17B-128E-Instruct-FP8 open source?

Llama-4-Maverick-17B-128E-Instruct-FP8 is an open-weight model. Check the license terms before commercial use.

### When was Llama-4-Maverick-17B-128E-Instruct-FP8 released?

Meta released Llama-4-Maverick-17B-128E-Instruct-FP8 on 5 Apr 2025, according to models.dev.

---
Data licensed CC BY 4.0. Cite as: AI Model Directory, "Llama-4-Maverick-17B-128E-Instruct-FP8", https://www.aimodel.directory/models/llama-4-maverick-17b-128e-instruct-fp8