Parameter counts are back on the menu, which means the AI industry is once again reviewing restaurants by counting the stoves. Reuters, carried by WMBD, reports that Thinking Machines Lab has revealed Inkling, its first general purpose model release from the San Francisco startup founded by former OpenAI CTO Mira Murati. The headline number is enormous, but the useful story is smaller and more practical: open weights, multimodal inputs, and a company openly positioning a model around customization rather than a confetti cannon of benchmark dominance. ## What happened, according to MarkTechPost and Databricks MarkTechPost describes Inkling as a 975 billion parameter open-weights multimodal mixture of experts model with 41 billion active parameters and controllable thinking effort. Mixture of Experts sounds like a consulting firm that charges by the espresso shot, but the core idea is straightforward: the model can be huge while activating only part of itself for a given computation. AlphaSignal reports that Inkling reasons over text, images, and audio, while Databricks says the model is available through Unity AI Gateway for building agents and applications on enterprise data. Databricks also says teams can govern Inkling with centralized security, permissions, cost controls, and observability, and connect it to coding agents such as Cursor and OpenCode. That distribution detail matters because open-weight releases do not win adoption by existing nobly on a download page, like a kayak in a garage. They need inference paths, governance hooks, and boring enterprise plumbing, which is where many technically good models go to discover paperwork. Inkling arriving inside existing developer and enterprise platforms is the part builders should circle in red pen. ## The benchmark caveat, according to VentureBeat VentureBeat reports that Inkling has high, but below state of the art, performance for open weights models on third party benchmarks. On SWE-bench Verified, VentureBeat says Inkling scores 77.6 percent, ahead of Nvidia Nemotron 3 at 71.9 percent. The same report says Inkling reaches 91.4 percent on VoiceBench. That is impressive, but it is not the usual model launch ritual where someone waves a bar chart around like it contains the secrets of civilization. This is the healthy part. A model can be valuable without being the single tallest benchmark skyscraper in a very foggy skyline. If you are building coding agents, voice workflows, or multimodal internal tools, the decision is not simply which model wins one public table. It is whether the model can be adapted, governed, deployed near private data, and made cheap enough that finance does not quietly replace your GPU budget with a scented candle. ## Why open weight is the strategy, according to Reuters via WMBD Reuters, carried by WMBD, explains that open-weight means users can download, run, and customize the underlying systems, unlike proprietary closed-source models. The same Reuters report says Inkling is available on Tinker and other developer platforms, and notes that Thinking Machines launched Tinker, a product for customizing AI models, last October. That sequence makes Inkling look less like a one-off model drop and more like a missing puzzle piece for a customization stack. First give builders tooling, then give them a large model worth tuning, which is suspiciously close to a plan. Reuters also frames Inkling as one of the few alternatives to popular open-source offerings from Chinese AI labs, while noting that the Western open-source ecosystem has lagged after Meta changed course toward a proprietary approach. VentureBeat adds another enterprise angle, reporting interest in open weights models that organizations can customize, control, and run on-premises or in virtual private clouds. Translation: some buyers want strong models, but they also want the right to inspect the machinery before inviting it into the vault. Rude of them to care about data control, really. ## What to watch next, according to AlphaSignal and VentureBeat AlphaSignal reports that Inkling was trained from scratch with full weights publicly available and trained to exhibit safe behavior across modalities. VentureBeat reports that Thinking Machines designed Inkling to answer directly on topics that may be subject to censorship, which will make evaluation more complicated than a leaderboard screenshot. Safety behavior, openness, and direct answering are not just marketing adjectives; they are deployment constraints with legal, policy, and product consequences. If your app touches enterprise data, regulated workflows, or user generated inputs, those constraints are where the rubber meets the audit log. For builders, the next question is not whether Inkling humiliates every closed model in a benchmark cage match. Watch how well it fine tunes, how much it costs to serve, how reliable its multimodal behavior is, and whether platform integrations make customization feel like engineering rather than archaeology. Thinking Machines Lab has made a credible first move by pairing a massive open-weight model with ecosystem access. The benchmark crown can stay in its glass case; the interesting part is whether people actually build with the thing. ## Sources - Thinking Machines Lab Releases Inkling: A 975B- ...
- Inkling model from Thinking Machines Lab now on ...
- Mira Murati's Thinking Machines Drops Inkling, a 975B Open ...
- Thinking Machines open sources first multimodal language ...
- AI startup Thinking Machines launches an open-weight AI model | 1470 & 100.3 WMBD
Sources
- Thinking Machines Lab Releases Inkling: A 975B- ...
- Mira Murati's Thinking Machines Lab has released 'Inkling ...
- Thinking Machines Lab Releases Inkling Open-Weights ...
- Inkling model from Thinking Machines Lab now on ...
- Together AI brings Thinking Machines Lab's new model ...
- Mira Murati's Thinking Machines Drops Inkling, a 975B Open ...
- MTS on X: "SITUATION DETECTED: Thinking Machines has released its first AI model, Inkling, an open-weight multimodal model built to be fine-tuned. Inkling is a 975B parameter MoE model with 41B active, a 1M token context window, trained on 45T tokens across text, image, audio, and video." / X
- Inkling model from Thinking Machines Lab now on ...
- Thinking Machines open sources first multimodal language ...
- AI startup Thinking Machines launches an open-weight AI model | 1470 & 100.3 WMBD