A gradual opening for open-weights models

Thinking Machines is testing a progressive release method with Inkling, an alternative to the immediate publication of model weights.

Thinking Machines proposes a step-by-step approach to making open-weights models accessible without treating their release as an automatic decision. The company distinguishes between two criteria: the dangerous capabilities inherent to the model and the ecosystem's capacity to detect, mitigate, or correct malicious uses.

For Inkling and Inkling-Small, the evaluation combined internal testing, analyses conducted by four external organizations, and fine-tuning designed to remove refusal behaviors. The work covered offensive cybersecurity, chemical and biological risks, interactions with vulnerable audiences, agentic uses, and concealment or sabotage behaviors.

According to Thinking Machines, these trials did not show any significant additional risk compared to already accessible open-weights models. The versions modified to respond to dangerous queries also reportedly did not exceed the capabilities observed in comparable models. This finding served as the basis for releasing the weights of Inkling and Inkling-Small.

For more advanced systems, the company envisions a progression ranging from limited API access to hosted fine-tuning, then to monitored public access before a potential release of the weights. Each step is intended to allow for the study of use cases, the preparation of defense teams, and the reassessment of risk levels before broadening access.

Thinking Machines is also exploring the possibility of reducing certain dangerous knowledge during training without affecting the model's general capabilities. This approach relies notably on filtering specialized documents, though the company specifies that this remains an active research question. Funding dedicated to security on the Tinker platform is also to be offered to the community.