Thinking Machines Lab released Inkling on July 15. The multimodal mixture-of-experts model has 975 billion total parameters, with 41 billion active per token, and supports a context window of up to one million tokens. The company says it pretrained Inkling on 45 trillion tokens spanning text, images, audio, and video.
| Specification | Inkling |
|---|---|
| Total parameters | 975B |
| Active parameters per token | 41B |
| Context window | Up to 1M tokens |
| Pretraining volume | 45T tokens |
| License | Apache 2.0 |
| Inputs | Text, images, audio |
The model card lists an Apache 2.0 license and links to downloadable weights on Hugging Face. Thinking Machines provides the original checkpoint and an NVFP4 checkpoint for Blackwell systems. The release also names support across SGLang, vLLM, TokenSpeed, llama.cpp, and Hugging Face Transformers, giving operators several deployment paths.
Thinking Machines describes Inkling as a base for customization and says plainly that it is "not the strongest overall model available today, open or closed." Fine-tuning is available through the company's Tinker service. The weights can also run through third-party inference providers or on infrastructure chosen by the operator.
Those routes carry different control boundaries. A self-hosted deployment can keep inference inside an organization's environment when configured correctly. Tinker and other hosted APIs send workloads through an external service. Apache 2.0 permits broad use and modification, while downstream libraries, model derivatives, and deployment components still require their own license review.
The operator test is straightforward: benchmark Inkling on the actual task, measure serving cost and latency, then measure the uplift from fine-tuning. The model card advises additional domain validation and human oversight for medical, legal, and safety-critical uses. Fine-tuning can also change safety behavior, which Thinking Machines identifies as an area it continues to study.
Independent task evaluations, deployment cost reports, and useful derivative checkpoints will show whether Inkling's customization pitch holds up outside the launch material.
The Signal is the public edge of a private practice. Sherpa points the same intelligence engine at one owner's business — competitors, suppliers, regulators, watched daily, graded and sourced. Work with a Sherpa →
