Thinking Machines Lab announced Inkling — its open-weights model — via @thinkymachines.
What it is
- 975B total parameters, 41B active — a Mixture-of-Experts transformer.
- Natively multimodal: text, images, audio, and video, over a 1M-token context window.
- Pretrained on 45 trillion tokens.
What it's built for
- Agentic coding with tool use.
- Controllable reasoning effort — dial the compute/quality trade-off per task.
- Strong multimodal performance and calibrated uncertainty (it expresses how sure it is).
- Positioned as "a practical multimodal foundation model for customization across domains" rather than a single-benchmark optimizer.
Availability
- Fine-tunable on Thinking Machines' Tinker platform.
- Inference via deployment partners: TogetherAI, Fireworks, Modal, and others.
- A smaller Inkling-Small (12B active) is previewed for latency-sensitive applications.
Why it matters
An open-weights, natively-multimodal MoE at this scale — with per-task reasoning-effort control and calibrated uncertainty — is a serious entry in the open frontier, and the Tinker + multi-provider story makes it something teams can actually customize and deploy rather than just benchmark.
Sources:
About the Authors
Federico Ulfo
Founder, Engineer
New York City