Llama-4-Maverick-17B-128E-Instruct-FP8
Summary
Llama-4-Maverick-17B-128E-Instruct-FP8 is an open-weight language model from Meta, released on 5 Apr 2025. It accepts text and images and generates text, with a 128K-token context window and up to 4K output tokens per response.
Open multimodal Llama model for strong reasoning and fast responses
Last verified 4 Oct 2026. Every figure on this page links to its source.
Specifications
| Developer | Meta |
|---|---|
| Release date | 5 Apr 2025 [models.dev] |
| Status | Generally available [models.dev] |
| Weights | Open weights [models.dev] |
| Context window | 128K (128,000 tokens) [models.dev] |
| Max output | 4K tokens [models.dev] |
| Knowledge cutoff | Aug 2024 [models.dev] |
| Input | Text and images [models.dev] |
| Output | Text [models.dev] |
| Reasoning mode | No [models.dev] |
| Tool calling | Yes [models.dev] |
Llama-4-Maverick-17B-128E-Instruct-FP8 API pricing by provider
USD per million tokens. Sorted by input price.
| Provider | Input | Output | Cached input | Context | Source |
|---|---|---|---|---|---|
| Llama API llama-4-maverick-17b-128e-instruct-fp8 | Free | Free | – | 128K | models.dev, 4 Oct 2026 |
Frequently asked questions
What is the context window of Llama-4-Maverick-17B-128E-Instruct-FP8?
Llama-4-Maverick-17B-128E-Instruct-FP8 has a context window of 128,000 tokens (128K) and can generate up to 4,096 tokens in a single response, according to models.dev.
Is Llama-4-Maverick-17B-128E-Instruct-FP8 open source?
Llama-4-Maverick-17B-128E-Instruct-FP8 is an open-weight model. Check the license terms before commercial use.
When was Llama-4-Maverick-17B-128E-Instruct-FP8 released?
Meta released Llama-4-Maverick-17B-128E-Instruct-FP8 on 5 Apr 2025, according to models.dev.
Change history
No changes detected since we started tracking this model. We re-check every source daily.