Inkling
Summary
Inkling is an open-weight language model from Thinkingmachines, released on 15 Jul 2026. It accepts text, images and audio and generates text, with a 524K-token context window and up to 1M output tokens per response. As of 4 Oct 2026, the lowest listed API price is $0.95 per million input tokens and $4.05 per million output tokens (DeepInfra), across 7 providers we track.
Multimodal MoE reasoning model (975B total, 41B active) for text, image, and audio
Last verified 4 Oct 2026. Every figure on this page links to its source.
Specifications
| Developer | Thinkingmachines |
|---|---|
| Release date | 15 Jul 2026 [models.dev] |
| Status | Generally available [models.dev] |
| Weights | Open weights [Hugging Face] |
| License | apache-2.0 [Hugging Face] |
| Parameters | 952B [Hugging Face] |
| Context window | 524K (524,288 tokens) [models.dev] |
| Max output | 1M tokens [models.dev] |
| Input | Text, images and audio [models.dev] |
| Output | Text [models.dev] |
| Reasoning mode | Yes [models.dev] |
| Tool calling | Yes [models.dev] |
| Structured output | Yes [models.dev] |
| Hugging Face | thinkingmachines/Inkling [Hugging Face] |
Inkling API pricing by provider
USD per million tokens. Sorted by input price.
| Provider | Input | Output | Cached input | Context | Source |
|---|---|---|---|---|---|
| DeepInfra thinkingmachines/Inkling | $0.95 | $4.05 | $0.16 | 524K | models.dev, 4 Oct 2026 |
| OpenRouter thinkingmachines/inkling | $0.95 | $4.05 | $0.16 | 524K | OpenRouter, 4 Oct 2026 |
| Hugging Face Inference Providers thinkingmachines/Inkling | $1 | $4.05 | – | 1M | models.dev, 4 Oct 2026 |
| Baseten thinkingmachines/inkling | $1 | $4.05 | – | 1M | models.dev, 4 Oct 2026 |
| Together AI thinkingmachines/Inkling | $1 | $4.05 | $0.17 | 524K | models.dev, 4 Oct 2026 |
| Fireworks AI accounts/fireworks/models/inkling | $1 | $4.05 | $0.17 | 1M | models.dev, 4 Oct 2026 |
| Vercel AI Gateway thinkingmachines/inkling | $1 | $4.05 | $0.17 | 256K | models.dev, 4 Oct 2026 |
Frequently asked questions
What is the context window of Inkling?
Inkling has a context window of 524,288 tokens (524K) and can generate up to 1,048,576 tokens in a single response, according to models.dev.
How much does the Inkling API cost?
Inkling is listed from $0.95 per million input tokens and $4.05 per million output tokens (DeepInfra). Prices last verified 4 Oct 2026.
Is Inkling open source?
Inkling is an open-weight model: its weights are published on Hugging Face as thinkingmachines/Inkling under the apache-2.0 license. Check the license terms before commercial use.
When was Inkling released?
Thinkingmachines released Inkling on 15 Jul 2026, according to models.dev.
Which providers offer Inkling?
We track Inkling on 7 providers: DeepInfra, OpenRouter, Hugging Face Inference Providers, Baseten, Together AI, Fireworks AI, Vercel AI Gateway.
Change history
No changes detected since we started tracking this model. We re-check every source daily.
Sources
- https://openrouter.ai/thinkingmachines/inkling
- https://github.com/sst/models.dev/blob/dev/providers/deepinfra/models/thinkingmachines/Inkling.toml
- https://huggingface.co/thinkingmachines/Inkling
- https://github.com/sst/models.dev/blob/dev/providers/huggingface/models/thinkingmachines/Inkling.toml
- https://github.com/sst/models.dev/blob/dev/providers/baseten/models/thinkingmachines/inkling.toml
- https://github.com/sst/models.dev/blob/dev/providers/togetherai/models/thinkingmachines/Inkling.toml
- https://github.com/sst/models.dev/blob/dev/providers/fireworks-ai/models/accounts/fireworks/models/inkling.toml
- https://github.com/sst/models.dev/blob/dev/providers/vercel/models/thinkingmachines/inkling.toml