Meta

Llama-4-Scout-17B-16E-Instruct-FP8

Generally availableOpen weights128K context

Summary

Llama-4-Scout-17B-16E-Instruct-FP8 is an open-weight language model from Meta, released on 5 Apr 2025. It accepts text and images and generates text, with a 128K-token context window and up to 4K output tokens per response.

Open multimodal Llama model for long-context analysis and efficient agents

Last verified 4 Oct 2026. Every figure on this page links to its source.

Specifications

DeveloperMeta
Release date5 Apr 2025 [models.dev]
StatusGenerally available [models.dev]
WeightsOpen weights [models.dev]
Context window128K (128,000 tokens) [models.dev]
Max output4K tokens [models.dev]
Knowledge cutoffAug 2024 [models.dev]
InputText and images [models.dev]
OutputText [models.dev]
Reasoning modeNo [models.dev]
Tool callingYes [models.dev]

Llama-4-Scout-17B-16E-Instruct-FP8 API pricing by provider

USD per million tokens. Sorted by input price.

ProviderInputOutputCached inputContextSource
Llama API
llama-4-scout-17b-16e-instruct-fp8
FreeFree–128Kmodels.dev, 4 Oct 2026

Frequently asked questions

What is the context window of Llama-4-Scout-17B-16E-Instruct-FP8?

Llama-4-Scout-17B-16E-Instruct-FP8 has a context window of 128,000 tokens (128K) and can generate up to 4,096 tokens in a single response, according to models.dev.

Is Llama-4-Scout-17B-16E-Instruct-FP8 open source?

Llama-4-Scout-17B-16E-Instruct-FP8 is an open-weight model. Check the license terms before commercial use.

When was Llama-4-Scout-17B-16E-Instruct-FP8 released?

Meta released Llama-4-Scout-17B-16E-Instruct-FP8 on 5 Apr 2025, according to models.dev.

Change history

No changes detected since we started tracking this model. We re-check every source daily.

Sources