Grok Voice STT 1.0
Summary
Grok Voice STT 1.0 is a proprietary language model from xAI, released on 4 Aug 2026. It accepts audio and generates text, with a 15K-token context window and up to 15K output tokens per response.
Grok Voice STT 1.0 is xAI's speech-to-text model. It supports transcription with word-level timestamps, optional speaker diarization, and multichannel audio.
Last verified 4 Oct 2026. Every figure on this page links to its source.
Specifications
| Developer | xAI |
|---|---|
| Release date | 4 Aug 2026 [models.dev] |
| Status | Generally available [models.dev] |
| Weights | Proprietary (API only) [models.dev] |
| Context window | 15K (15,000 tokens) [models.dev] |
| Max output | 15K tokens [models.dev] |
| Input | Audio [models.dev] |
| Output | Text [models.dev] |
| Reasoning mode | No [models.dev] |
| Tool calling | No [models.dev] |
Grok Voice STT 1.0 API pricing by provider
USD per million tokens. Sorted by input price.
We do not currently track a paid API listing for this model.
Frequently asked questions
What is the context window of Grok Voice STT 1.0?
Grok Voice STT 1.0 has a context window of 15,000 tokens (15K) and can generate up to 15,000 tokens in a single response, according to models.dev.
Is Grok Voice STT 1.0 open source?
No. Grok Voice STT 1.0 is a proprietary model; xAI has not released its weights.
When was Grok Voice STT 1.0 released?
xAI released Grok Voice STT 1.0 on 4 Aug 2026, according to models.dev.
Change history
No changes detected since we started tracking this model. We re-check every source daily.