# Llama-3.3-70B-Instruct

> Llama-3.3-70B-Instruct is an open-weight language model from Meta, released on 6 Dec 2024. It accepts text and generates text, with a 128K-token context window and up to 4K output tokens per response. As of 4 Oct 2026, the lowest listed API price is $0.14 per million input tokens and $0.40 per million output tokens (Novita AI), across 8 providers we track.

Source page: https://www.aimodel.directory/models/llama-3-3-70b-instruct
Last verified: 4 Oct 2026

## Specifications

- **Developer:** Meta
- **Release date:** 6 Dec 2024 (source: models.dev, https://github.com/sst/models.dev/blob/dev/providers/llama/models/llama-3.3-70b-instruct.toml)
- **Status:** Generally available (source: models.dev, https://github.com/sst/models.dev/blob/dev/providers/llama/models/llama-3.3-70b-instruct.toml)
- **Weights:** Open weights (source: Hugging Face, https://huggingface.co/meta-llama/Llama-3.3-70B-Instruct)
- **License:** llama3.3 (source: Hugging Face, https://huggingface.co/meta-llama/Llama-3.3-70B-Instruct)
- **Parameters:** 70.6B (source: Hugging Face, https://huggingface.co/meta-llama/Llama-3.3-70B-Instruct)
- **Context window:** 128K (128,000 tokens) (source: models.dev, https://github.com/sst/models.dev/blob/dev/providers/llama/models/llama-3.3-70b-instruct.toml)
- **Max output:** 4K tokens (source: models.dev, https://github.com/sst/models.dev/blob/dev/providers/llama/models/llama-3.3-70b-instruct.toml)
- **Knowledge cutoff:** Dec 2023 (source: models.dev, https://github.com/sst/models.dev/blob/dev/providers/llama/models/llama-3.3-70b-instruct.toml)
- **Input:** Text (source: models.dev, https://github.com/sst/models.dev/blob/dev/providers/llama/models/llama-3.3-70b-instruct.toml)
- **Output:** Text (source: models.dev, https://github.com/sst/models.dev/blob/dev/providers/llama/models/llama-3.3-70b-instruct.toml)
- **Reasoning mode:** No (source: models.dev, https://github.com/sst/models.dev/blob/dev/providers/llama/models/llama-3.3-70b-instruct.toml)
- **Tool calling:** Yes (source: models.dev, https://github.com/sst/models.dev/blob/dev/providers/llama/models/llama-3.3-70b-instruct.toml)
- **Hugging Face:** meta-llama/Llama-3.3-70B-Instruct (source: Hugging Face, https://huggingface.co/meta-llama/Llama-3.3-70B-Instruct)

## API pricing (USD per 1M tokens)

| Provider | Input | Output | Cached input | Context |
|---|---|---|---|---|
| Llama API | Free | Free | – | 128K |
| Novita AI | $0.14 | $0.40 | – | 131K |
| OpenRouter | $0.22 | $0.50 | $0.11 | 131K |
| Cloudflare Workers AI | $0.29 | $2.25 | – | 24K |
| Hugging Face Inference Providers | $0.59 | $0.79 | – | 131K |
| Azure AI Foundry | $0.71 | $0.71 | – | 128K |
| Amazon Bedrock | $0.72 | $0.72 | – | 128K |
| Amazon Bedrock | $0.72 | $0.72 | – | 128K |
| Snowflake Cortex | – | – | – | 128K |

## FAQ

### What is the context window of Llama-3.3-70B-Instruct?

Llama-3.3-70B-Instruct has a context window of 128,000 tokens (128K) and can generate up to 4,096 tokens in a single response, according to models.dev.

### How much does the Llama-3.3-70B-Instruct API cost?

Through Llama API, Llama-3.3-70B-Instruct costs Free per million input tokens and Free per million output tokens. Prices last verified 4 Oct 2026.

### Is Llama-3.3-70B-Instruct open source?

Llama-3.3-70B-Instruct is an open-weight model: its weights are published on Hugging Face as meta-llama/Llama-3.3-70B-Instruct under the llama3.3 license. Check the license terms before commercial use.

### When was Llama-3.3-70B-Instruct released?

Meta released Llama-3.3-70B-Instruct on 6 Dec 2024, according to models.dev.

### Which providers offer Llama-3.3-70B-Instruct?

We track Llama-3.3-70B-Instruct on 8 providers: Llama API, Novita AI, OpenRouter, Cloudflare Workers AI, Hugging Face Inference Providers, Azure AI Foundry, Amazon Bedrock, Snowflake Cortex.

---
Data licensed CC BY 4.0. Cite as: AI Model Directory, "Llama-3.3-70B-Instruct", https://www.aimodel.directory/models/llama-3-3-70b-instruct