# Inkling

> Inkling is an open-weight language model from Thinkingmachines, released on 15 Jul 2026. It accepts text, images and audio and generates text, with a 524K-token context window and up to 1M output tokens per response. As of 4 Oct 2026, the lowest listed API price is $0.95 per million input tokens and $4.05 per million output tokens (DeepInfra), across 7 providers we track.

Source page: https://www.aimodel.directory/models/inkling
Last verified: 4 Oct 2026

## Specifications

- **Developer:** Thinkingmachines
- **Release date:** 15 Jul 2026 (source: models.dev, https://github.com/sst/models.dev/blob/dev/providers/deepinfra/models/thinkingmachines/Inkling.toml)
- **Status:** Generally available (source: models.dev, https://github.com/sst/models.dev/blob/dev/providers/deepinfra/models/thinkingmachines/Inkling.toml)
- **Weights:** Open weights (source: Hugging Face, https://huggingface.co/thinkingmachines/Inkling)
- **License:** apache-2.0 (source: Hugging Face, https://huggingface.co/thinkingmachines/Inkling)
- **Parameters:** 952B (source: Hugging Face, https://huggingface.co/thinkingmachines/Inkling)
- **Context window:** 524K (524,288 tokens) (source: models.dev, https://github.com/sst/models.dev/blob/dev/providers/deepinfra/models/thinkingmachines/Inkling.toml)
- **Max output:** 1M tokens (source: models.dev, https://github.com/sst/models.dev/blob/dev/providers/deepinfra/models/thinkingmachines/Inkling.toml)
- **Input:** Text, images and audio (source: models.dev, https://github.com/sst/models.dev/blob/dev/providers/deepinfra/models/thinkingmachines/Inkling.toml)
- **Output:** Text (source: models.dev, https://github.com/sst/models.dev/blob/dev/providers/deepinfra/models/thinkingmachines/Inkling.toml)
- **Reasoning mode:** Yes (source: models.dev, https://github.com/sst/models.dev/blob/dev/providers/deepinfra/models/thinkingmachines/Inkling.toml)
- **Tool calling:** Yes (source: models.dev, https://github.com/sst/models.dev/blob/dev/providers/deepinfra/models/thinkingmachines/Inkling.toml)
- **Structured output:** Yes (source: models.dev, https://github.com/sst/models.dev/blob/dev/providers/deepinfra/models/thinkingmachines/Inkling.toml)
- **Hugging Face:** thinkingmachines/Inkling (source: Hugging Face, https://huggingface.co/thinkingmachines/Inkling)

## API pricing (USD per 1M tokens)

| Provider | Input | Output | Cached input | Context |
|---|---|---|---|---|
| DeepInfra | $0.95 | $4.05 | $0.16 | 524K |
| OpenRouter | $0.95 | $4.05 | $0.16 | 524K |
| Hugging Face Inference Providers | $1 | $4.05 | – | 1M |
| Baseten | $1 | $4.05 | – | 1M |
| Together AI | $1 | $4.05 | $0.17 | 524K |
| Fireworks AI | $1 | $4.05 | $0.17 | 1M |
| Vercel AI Gateway | $1 | $4.05 | $0.17 | 256K |

## FAQ

### What is the context window of Inkling?

Inkling has a context window of 524,288 tokens (524K) and can generate up to 1,048,576 tokens in a single response, according to models.dev.

### How much does the Inkling API cost?

Inkling is listed from $0.95 per million input tokens and $4.05 per million output tokens (DeepInfra). Prices last verified 4 Oct 2026.

### Is Inkling open source?

Inkling is an open-weight model: its weights are published on Hugging Face as thinkingmachines/Inkling under the apache-2.0 license. Check the license terms before commercial use.

### When was Inkling released?

Thinkingmachines released Inkling on 15 Jul 2026, according to models.dev.

### Which providers offer Inkling?

We track Inkling on 7 providers: DeepInfra, OpenRouter, Hugging Face Inference Providers, Baseten, Together AI, Fireworks AI, Vercel AI Gateway.

---
Data licensed CC BY 4.0. Cite as: AI Model Directory, "Inkling", https://www.aimodel.directory/models/inkling