Skip to main content
AI Socratic

Thinking Machines: Introducing Inkling

Federico UlfoFederico Ulfo
July 22, 20262 min read

Thinking Machines Lab announced Inkling — its open-weights model — via @thinkymachines.

What it is

  • 975B total parameters, 41B active — a Mixture-of-Experts transformer.
  • Natively multimodal: text, images, audio, and video, over a 1M-token context window.
  • Pretrained on 45 trillion tokens.

What it's built for

  • Agentic coding with tool use.
  • Controllable reasoning effort — dial the compute/quality trade-off per task.
  • Strong multimodal performance and calibrated uncertainty (it expresses how sure it is).
  • Positioned as "a practical multimodal foundation model for customization across domains" rather than a single-benchmark optimizer.

Availability

  • Fine-tunable on Thinking Machines' Tinker platform.
  • Inference via deployment partners: TogetherAI, Fireworks, Modal, and others.
  • A smaller Inkling-Small (12B active) is previewed for latency-sensitive applications.

Why it matters

An open-weights, natively-multimodal MoE at this scale — with per-task reasoning-effort control and calibrated uncertainty — is a serious entry in the open frontier, and the Tinker + multi-provider story makes it something teams can actually customize and deploy rather than just benchmark.

Sources:

About the Authors

Federico Ulfo

Federico Ulfo

Founder, Engineer

New York City