OpenAI

GPT Audio

Generally available128K context

Summary

GPT Audio is a language model from OpenAI, released on 19 Jan 2026. It accepts text and audio and generates text and audio, with a 128K-token context window and up to 16K output tokens per response. As of 4 Oct 2026, the lowest listed API price is $2.50 per million input tokens and $10 per million output tokens (OpenRouter), across 1 provider we track.

The gpt-audio model is OpenAI's first generally available audio model. The new snapshot features an upgraded decoder for more natural sounding voices and maintains better voice consistency. Audio is priced...

Last verified 4 Oct 2026. Every figure on this page links to its source.

Specifications

DeveloperOpenAI
Release date19 Jan 2026 [OpenRouter]
StatusGenerally available
Context window128K (128,000 tokens) [OpenRouter]
Max output16K tokens [OpenRouter]
InputText and audio [OpenRouter]
OutputText and audio [OpenRouter]
Reasoning modeNo [OpenRouter]
Tool callingYes [OpenRouter]
Structured outputYes [OpenRouter]

GPT Audio API pricing by provider

USD per million tokens. Sorted by input price.

ProviderInputOutputCached inputContextSource
OpenRouter
openai/gpt-audio
$2.50$10–128KOpenRouter, 4 Oct 2026

Frequently asked questions

What is the context window of GPT Audio?

GPT Audio has a context window of 128,000 tokens (128K) and can generate up to 16,384 tokens in a single response, according to OpenRouter.

How much does the GPT Audio API cost?

GPT Audio is listed from $2.50 per million input tokens and $10 per million output tokens (OpenRouter). Prices last verified 4 Oct 2026.

When was GPT Audio released?

OpenAI released GPT Audio on 19 Jan 2026, according to OpenRouter.

Change history

No changes detected since we started tracking this model. We re-check every source daily.

Sources