Policypublished

Thinking Machines Releases Inkling, but Argues Open Weights Need a Gate

The lab’s argument is not that every capable model should be closed. It is that release decisions should turn on whether a model adds risk beyond what is already downloadable—and whether defenders have had time to prepare.

By 3 min read
Thinking Machines Releases Inkling, but Argues Open Weights Need a Gate

Listen to this story

The audio brief

About 1:29
0:001:29
Read transcript
Thinking Machines Lab has released Inkling and Inkling-Small as downloadable open-weight models, while arguing that open weights should usually be the final step in a safety process—not the first. The lab’s proposed alternative starts with monitored access through an API, then hosted fine-tuning, monitored general availability, and only eventually public weights. The idea is to let users and defenders work with increasingly capable systems while providers can still monitor use, maintain safeguards, and revoke access. For Inkling, the question was not whether the models were harmless. It was whether they added material risk beyond comparable models that are already downloadable. Thinking Machines says its assessment covered chemical, biological, radiological, and nuclear, or CBRN, risks; offensive cyber use; agentic misuse; loss of control; and harmful multimodal scenarios, across 17 languages and text, image, and audio inputs. External testing came from Scale AI, Handshake AI, FAR.AI, and Apollo Research. The lab also adversarially fine-tuned variants to remove refusal behavior. It says those variants showed no new CBRN or cyber capability uplift beyond the existing open-weight baseline. But the framework is still only a high-level proposal. The unresolved issue is operational: what evidence should move a model to the next access stage, and how should defensive readiness be measured?

Story brief

3 key points

Thinking Machines Lab has made Inkling and Inkling-Small downloadable while using the release to propose a conditional alternative to immediate public weights. The lab’s framework would move models through monitored APIs, hosted fine-tuning, monitored general availability and only then open weights, with each step tied to evidence and defensive readiness. Its Inkling assessment—covering 17 languages, multiple...

  1. 01

    Inkling’s evaluation covered CBRN, offensive cyber, agentic misuse, loss of control and harmful multimodal scenarios.

  2. 02

    Scale AI, Handshake AI, FAR.AI and Apollo Research conducted external testing across distinct risk areas.

  3. 03

    Adversarial fine-tuning removed refusal behavior but produced no new CBRN or cyber capability uplift, according to Thinking Machines.

Thinking Machines Lab has released Inkling and Inkling-Small as open-weight language models, while arguing that public weights should be the last step of a safety process rather than the default starting point. Its proposed middle path is to widen access only as evidence supports it and as the surrounding defensive ecosystem becomes more capable.

Publication versus progression

The contrast is between two release logics. Publishing weights lets users own, run and customize a model, but it also makes access irreversible and enables modifications that can remove safety training. Thinking Machines says open models can distribute development and safety work beyond a small group of labs, yet the same distribution creates real misuse risk.

Its alternative is not a fixed ladder that automatically ends in public weights. The lab proposes monitored inference API access, hosted fine-tuning, monitored general availability and, eventually, open weights. Each expansion is meant to give defenders access to capable systems while preserving a provider’s ability to monitor use, maintain guardrails and revoke access before weights are released.

Why Inkling cleared a different bar

For Inkling, the lab did not claim the models were harmless. It framed the decision more narrowly: whether their release would add material incremental risk beyond existing open-weight models. That is a comparative standard, shaped by the lab’s assessment that models of comparable or greater strength in relevant dangerous domains are already downloadable.

The release assessment had three layers

  • Internal tests covered CBRN topics, offensive cybersecurity, broad misuse in agentic and tool-use settings, and multimodal harmful-content scenarios across 17 languages and text, image and audio inputs.
  • Four outside organizations tested distinct areas: Scale AI for general misuse, Handshake AI for vulnerable-user interactions, FAR.AI for CBRN and cybersecurity, and Apollo Research for loss-of-control behaviors.
  • The lab also fine-tuned variants to comply with harmful requests, testing the capabilities exposed after refusal behavior was removed.

Thinking Machines says those tests did not find capabilities that would materially raise real-world risk beyond the existing open-weight baseline. Its helpful-only, adversarially fine-tuned variants produced no new uplift on CBRN or cyber evaluations and remained comparable to existing open-weight models, according to the lab.

A framework designed for a harder future case

The Inkling conclusion does not settle how the framework would work for models nearer the capability frontier. The lab says the balance will change as its systems become more capable, more accessible or easier to modify. Its proposed response is to test dangerous capabilities directly and to build layered defenses before opening access further.

There is a second, more ambitious part of the argument: dangerous specialized knowledge may be separable from general intelligence. Thinking Machines points to pretraining-data curation and post-training interventions as possible ways to reduce dangerous capabilities without broadly degrading useful ones. It treats that possibility as an active research question, not a proven safeguard; a sufficiently capable model might re-derive filtered knowledge, and the boundary between dangerous and ordinary technical knowledge may be unclear.

The decisive unanswered question is operational: what evidence should move a model from one access stage to the next, and how should ecosystem readiness be measured? Thinking Machines calls its post a high-level framework and says it plans a more detailed version covering evaluations and access criteria. Until those thresholds are specified, the proposal is a direction for release governance rather than a release standard others can independently apply.

Sources

  1. thinkingmachines.aiA Safe Path to Open Weights