Google DeepMind

Gemma 4 12B IT

Generally availableProprietary262K context

Summary

Gemma 4 12B IT is a proprietary language model from Google DeepMind, released on 9 Jun 2026. It accepts text and generates text, with a 262K-token context window and up to 262K output tokens per response. As of 4 Oct 2026, the lowest listed API price is $0.10 per million input tokens and $0.30 per million output tokens (SiliconFlow), across 1 provider we track.

Compact Gemma 4 instruction model for open, self-hosted chat and reasoning

Last verified 4 Oct 2026. Every figure on this page links to its source.

Specifications

DeveloperGoogle DeepMind
Release date9 Jun 2026 [models.dev]
StatusGenerally available [models.dev]
WeightsProprietary (API only) [models.dev]
Context window262K (262,144 tokens) [models.dev]
Max output262K tokens [models.dev]
InputText [models.dev]
OutputText [models.dev]
Reasoning modeNo [models.dev]
Tool callingYes [models.dev]
Structured outputYes [models.dev]

Gemma 4 12B IT API pricing by provider

USD per million tokens. Sorted by input price.

ProviderInputOutputCached inputContextSource
SiliconFlow
google/gemma-4-12B-it
$0.10$0.30–262Kmodels.dev, 4 Oct 2026

Frequently asked questions

What is the context window of Gemma 4 12B IT?

Gemma 4 12B IT has a context window of 262,144 tokens (262K) and can generate up to 262,144 tokens in a single response, according to models.dev.

How much does the Gemma 4 12B IT API cost?

Gemma 4 12B IT is listed from $0.10 per million input tokens and $0.30 per million output tokens (SiliconFlow). Prices last verified 4 Oct 2026.

Is Gemma 4 12B IT open source?

No. Gemma 4 12B IT is a proprietary model; Google DeepMind has not released its weights. It is available through APIs including SiliconFlow.

When was Gemma 4 12B IT released?

Google DeepMind released Gemma 4 12B IT on 9 Jun 2026, according to models.dev.

Change history

No changes detected since we started tracking this model. We re-check every source daily.

Sources