Liquid AI releases two models capable of making decisions without generating text
Liquid AI's d1 models eliminate token generation to accelerate classification and routing—an approach that reduces local latency.
Liquid AI releases d1-3B and d1-omni-600M, two open-weight models designed not to draft a response, but to make a structured decision in a single pass. The d1 family targets use cases such as reranking, moderation, classification, agent guardrails, visual inspection, or voice command routing.
This principle sets them apart from classic generative models. The d1 models do not produce any output tokens. They receive a state—for example, text, JSON, an image, or an audio signal—along with a series of questions or options, and then directly return typed answers from their probability distribution. This architecture mechanically reduces the time spent generating and parsing a text response.
The largest model, d1-3B available on Hugging Face, is based on LFM2.5-VL-3B and simultaneously accepts text and images. It has approximately 3.1 billion parameters and features a context window of 32,768 tokens. Liquid AI recommends it in particular for routing, classification, response evaluation, agent guardrails, and visual inspection.
According to evaluations published by Liquid AI, d1-3B scores 48.57 on the Decision Index v0.2.1, placing it at the top of the models under 10 billion parameters tested on this benchmark, and slightly ahead of Decider 35B-A3B, which is credited with 47.11. Across seven public text benchmarks, Liquid AI also reports an average of 82.9, ahead of Decider 4B at 81.1. These results remain those communicated by the company.
The second model, d1-omni-600M, is much more compact. Its actual size is about 587 million parameters, and it is based on LFM2.5-Encoder-350M, to which Liquid AI added dedicated vision and audio encoders. It can process text + image or text + audio, with up to 30 seconds of speech in a single pass.
This experimental variant primarily targets devices where the memory footprint becomes critical. Liquid AI cites voice command routing, local moderation, or intent classification. In its own text comparisons, d1-omni-600M notably achieves 95.8 on Civil Comments and 79.5 on PAWS-X, the highest scores in the table published by the company for toxicity detection and paraphrase identification.
The other focus of the launch directly concerns local inference. Liquid AI announces out-of-the-box support for llama.cpp and operation across the entire NVIDIA lineup, from DGX to RTX machines and embedded Jetson devices. On d1-3B, the company measures 8 ms for a simple decision on an RTX 4090, 16 ms on Jetson AGX Thor, 26 ms on AGX Orin 64 GB, and 50 ms on Orin Nano. On AMD MI325X, the same test drops to 9 ms.
For visual use cases, Liquid AI measures the processing of a 384 px image at 17 ms on RTX 4090, 18 ms on AMD MI325X, 35 ms on Jetson AGX Thor, and 202 ms on Orin Nano. These figures correspond to measurements taken by Liquid AI, one request at a time, and therefore do not constitute independent real-world deployment benchmarks.
Liquid AI also provides a quantized version of d1-3B, intended notably for NVIDIA GPUs and Jetsons. d1-3B-w8a8 uses INT8 weights and activations, reducing the checkpoint size from approximately 6.2 GB to 3.8 GB. The goal is to leave more memory available for other components of an embedded application.
To illustrate these use cases, Liquid AI built ten demonstrations based on a live camera feed, ranging from gesture-controlled gaming to visual moderation. Each frame is evaluated through a single pass of the model. In collaboration with NVIDIA, the company also shows d1-3B driving navigation in Isaac Sim with the model running on a Jetson.
The d1 family does not come out of nowhere. A few days earlier, Liquid AI had presented its first proprietary d1 model, capable of analyzing text and images and compared to GPT-6.1 Sol and Claude Opus 5.5 across several decision-making scenarios. With Open d1, the company is now bringing this approach to downloadable, locally executable models.
The value of d1 therefore lies less in its ability to converse than in its very refusal to do so. Where a general-purpose model must generate a response before a system can interpret it, d1 is designed to directly return an actionable decision. On a server, a workstation, or an embedded device, this specialization can sharply reduce latency and cost for any task where a long response adds no value.