# Gemini 2.5 Flash-Lite

> Gemini 2.5 Flash-Lite is a proprietary language model from Google DeepMind, released on 17 Jun 2025. It accepts text, images, audio, video and PDFs and generates text, with a 1M-token context window and up to 66K output tokens per response. As of 4 Oct 2026, the lowest listed API price is $0.10 per million input tokens and $0.40 per million output tokens (Gemini API), across 4 providers we track. It is scheduled for retirement on 20 Oct 2026.

Source page: https://www.aimodel.directory/models/gemini-2-5-flash-lite
Last verified: 4 Oct 2026

## Specifications

- **Developer:** Google DeepMind
- **Release date:** 17 Jun 2025 (source: models.dev, https://github.com/sst/models.dev/blob/dev/providers/google/models/gemini-2.5-flash-lite.toml)
- **Status:** Generally available (source: models.dev, https://github.com/sst/models.dev/blob/dev/providers/google/models/gemini-2.5-flash-lite.toml)
- **Weights:** Proprietary (API only) (source: models.dev, https://github.com/sst/models.dev/blob/dev/providers/google/models/gemini-2.5-flash-lite.toml)
- **Context window:** 1M (1,048,576 tokens) (source: models.dev, https://github.com/sst/models.dev/blob/dev/providers/google/models/gemini-2.5-flash-lite.toml)
- **Max output:** 66K tokens (source: models.dev, https://github.com/sst/models.dev/blob/dev/providers/google/models/gemini-2.5-flash-lite.toml)
- **Knowledge cutoff:** Jan 2025 (source: models.dev, https://github.com/sst/models.dev/blob/dev/providers/google/models/gemini-2.5-flash-lite.toml)
- **Input:** Text, images, audio, video and PDFs (source: models.dev, https://github.com/sst/models.dev/blob/dev/providers/google/models/gemini-2.5-flash-lite.toml)
- **Output:** Text (source: models.dev, https://github.com/sst/models.dev/blob/dev/providers/google/models/gemini-2.5-flash-lite.toml)
- **Reasoning mode:** Yes (source: models.dev, https://github.com/sst/models.dev/blob/dev/providers/google/models/gemini-2.5-flash-lite.toml)
- **Tool calling:** Yes (source: models.dev, https://github.com/sst/models.dev/blob/dev/providers/google/models/gemini-2.5-flash-lite.toml)
- **Structured output:** Yes (source: models.dev, https://github.com/sst/models.dev/blob/dev/providers/google/models/gemini-2.5-flash-lite.toml)
- **Retirement date:** 20 Oct 2026 (source: OpenRouter, https://openrouter.ai/google/gemini-2.5-flash-lite)

## API pricing (USD per 1M tokens)

| Provider | Input | Output | Cached input | Context |
|---|---|---|---|---|
| Gemini API | $0.10 | $0.40 | $0.01 | 1M |
| Google Vertex AI | $0.10 | $0.40 | $0.01 | 1M |
| OpenRouter | $0.10 | $0.40 | $0.01 | 1M |
| Vercel AI Gateway | $0.10 | $0.40 | $0.01 | 1M |

## FAQ

### What is the context window of Gemini 2.5 Flash-Lite?

Gemini 2.5 Flash-Lite has a context window of 1,048,576 tokens (1M) and can generate up to 65,536 tokens in a single response, according to models.dev.

### How much does the Gemini 2.5 Flash-Lite API cost?

Through Gemini API, Gemini 2.5 Flash-Lite costs $0.10 per million input tokens and $0.40 per million output tokens. Prices last verified 4 Oct 2026.

### Is Gemini 2.5 Flash-Lite open source?

No. Gemini 2.5 Flash-Lite is a proprietary model; Google DeepMind has not released its weights. It is available through APIs including Gemini API, Google Vertex AI, OpenRouter, Vercel AI Gateway.

### When was Gemini 2.5 Flash-Lite released?

Google DeepMind released Gemini 2.5 Flash-Lite on 17 Jun 2025, according to models.dev.

### Which providers offer Gemini 2.5 Flash-Lite?

We track Gemini 2.5 Flash-Lite on 4 providers: Gemini API, Google Vertex AI, OpenRouter, Vercel AI Gateway.

---
Data licensed CC BY 4.0. Cite as: AI Model Directory, "Gemini 2.5 Flash-Lite", https://www.aimodel.directory/models/gemini-2-5-flash-lite