xAI

Grok Voice STT 1.0

Generally availableProprietary15K context

Summary

Grok Voice STT 1.0 is a proprietary language model from xAI, released on 4 Aug 2026. It accepts audio and generates text, with a 15K-token context window and up to 15K output tokens per response.

Grok Voice STT 1.0 is xAI's speech-to-text model. It supports transcription with word-level timestamps, optional speaker diarization, and multichannel audio.

Last verified 4 Oct 2026. Every figure on this page links to its source.

Specifications

DeveloperxAI
Release date4 Aug 2026 [models.dev]
StatusGenerally available [models.dev]
WeightsProprietary (API only) [models.dev]
Context window15K (15,000 tokens) [models.dev]
Max output15K tokens [models.dev]
InputAudio [models.dev]
OutputText [models.dev]
Reasoning modeNo [models.dev]
Tool callingNo [models.dev]

Grok Voice STT 1.0 API pricing by provider

USD per million tokens. Sorted by input price.

We do not currently track a paid API listing for this model.

Frequently asked questions

What is the context window of Grok Voice STT 1.0?

Grok Voice STT 1.0 has a context window of 15,000 tokens (15K) and can generate up to 15,000 tokens in a single response, according to models.dev.

Is Grok Voice STT 1.0 open source?

No. Grok Voice STT 1.0 is a proprietary model; xAI has not released its weights.

When was Grok Voice STT 1.0 released?

xAI released Grok Voice STT 1.0 on 4 Aug 2026, according to models.dev.

Change history

No changes detected since we started tracking this model. We re-check every source daily.

Sources