Thinking Machines opened Inkling. Your server rack stayed closed.

Thinking Machines opened Inkling. Your server rack stayed closed.

Inkling's open weights are real, but running the 975B multimodal model starts with a GPU-cluster problem most developers will solve by renting it.

"Open weights. Closed reality."
Thinking Machines Lab released Inkling on July 15 as a model you can download, modify, and fine-tune. It also released a model whose full-precision checkpoint needs at least 2 TB of aggregated GPU memory. The gap between those two sentences is the entire product review. 1 2

What Inkling actually is

Inkling is a sparse Mixture-of-Experts transformer with 975 billion total parameters and about 41 billion active for each token. It accepts text, images, and audio, produces text, supports up to a 1 million-token context window, and was pretrained on 45 trillion tokens spanning text, images, audio, and video. Thinking Machines is not claiming the overall crown. It describes Inkling as a broad base for customization, with a controllable "thinking effort" setting that lets developers trade response depth against cost and latency. 1 2
That is a sensible pitch. A model does not need to win every benchmark if a company can adapt it to a narrow workflow. Inkling is aimed at developers building coding assistants, agent and tool-use systems, chatbots, and retrieval-augmented applications, not at someone looking for a chatbot to summarize a PDF on a laptop. 2

The word "open" is doing heavy lifting

The weights are available on Hugging Face under the Apache 2.0 license. That is real access, not a waitlist dressed up as openness. The hardware bill is also real:
RouteWhat the documentation saysWho it is really for
Self-hosted BF16At least 2 TB of aggregated VRAM, such as 8 NVIDIA B300 GPUs or 16 NVIDIA H200 GPUsTeams with a data-center-sized budget and people who can operate it 2
Self-hosted NVFP4At least 600 GB of aggregated VRAM, with supported Blackwell or H200 configurationsTeams that can optimize around a specific accelerator stack 2
Hosted or fine-tunedTinker for fine-tuning and playground access, plus listed inference partners including Together AI, Fireworks, Modal, Databricks, and BasetenDevelopers who want the model without owning the rack 1
This is the part that gets flattened into "download and run." You can download Inkling. Running it is a separate negotiation with memory bandwidth, orchestration software, power, and somebody who knows why the serving stack has stopped talking to the GPUs. The model card lists SGLang, vLLM, TokenSpeed, Unsloth, and Hugging Face among the supported software routes. Open weights remove one gate. They do not remove the building. 2

The demo proves the platform, not the discount

Thinking Machines' launch showcase has Inkling write its own fine-tuning job, create synthetic data and an evaluation, run the job through Tinker, and load the resulting checkpoint into an OpenCode harness. The company says that loop took about 27 minutes. It is a neat demonstration of what a customization platform can do when the platform, model, tools, and demo environment are already aligned. 1
It is not the same thing as making customization cheap or simple for everyone else. TechCrunch reports that Thinking Machines is positioning Inkling for enterprise customization and that doing serious fine-tuning still requires machine-learning expertise. The launch lists a temporary free Playground and a temporary 50% Tinker discount, but it does not present one simple Inkling subscription price. Hosted access is split across Tinker and third-party providers; self-hosting is an infrastructure project. 3 1
So the customer is not buying a cheap model. The customer is choosing where to pay: a hosted meter, a cluster, or a team of specialists. The download is the least expensive-looking part because it is the part the vendor can give away.

Open weights, downstream responsibility

Thinking Machines' Model Acceptable Use Policy applies when users access, download, or use the model materials. It requires legal use, puts responsibility for products built on the materials on the downstream user, and requires customer-facing deployments to disclose material limitations or risks and, where required, that users are interacting with an AI system. 4
That is not an unreasonable policy. A customizable model needs someone to own the consequences. But it changes the meaning of the sales pitch. The customer gets control over the weights and inherits a large part of the safety and deployment job. The model is open; the responsibility is even more open.

Verdict

Inkling is a credible customization base with unusually broad inputs and a genuinely permissive weight-release story. It is also a 975B model whose practical self-hosting path starts at a GPU cluster, not a developer workstation. Buy into it if you have a real reason to control the model, a provider budget, and the ML staff to tune and govern it. If you just want a capable multimodal model, the realistic product is hosted access. In that case, "open weights" is useful insurance for a future you may never have the hardware to exercise.

관련 콘텐츠

  • 로그인하면 댓글을 작성할 수 있습니다.
More from this channel