Open source model inference splits between prefill and generation at Prime Intellect
Prime Intellect is launching Prime Inference, its inference infrastructure for open-source models, featuring serverless endpoints and reserved capacity. For GLM-4 on NVIDIA GB200 NVL72, the company separates prefill from decode and optimizes the KV cache to maintain low-latency agents over long contexts.
Infrastructure Born from Prime’s Own Workloads
Nearly one trillion tokens are already processed daily internally by the infrastructure behind Prime Inference. It was used for reinforcement learning rollouts, synthetic data generation, evaluations, and coding agents before being opened as an inference service.
Prime Intellect also reports serving production customer deployments since January. Prime Inference now offers serverless endpoints for variable workloads and reserved capacity for continuous use, distributed across multiple datacenters.
The endpoints remain compatible with the OpenAI API. The current infrastructure uses NVIDIA Blackwell GPUs, with Vera Rubin on Prime’s roadmap. Separating Prompt Processing from Generation
Agents maintaining long conversations pose a particular challenge. In the workload studied by Prime, a single turn adds about 6,000 tokens to a prompt that can already reach 140,000 tokens, a large portion of which has already been computed during previous exchanges.
On shared GPUs, processing a new long prompt can slow down generation for already active sessions. Prime therefore separates prefill, dedicated to processing context, and decode, responsible for generating new tokens, onto distinct GPU groups.
NVIDIA Dynamo orchestrates and routes requests while vLLM runs the model. Once the prefill is complete, the computed KV cache is transferred to the decode group via NIXL. In Prime’s tests, this separation reduces p90 inter-token latency by nearly 40%. Retaining Rather Than Recomputing Context
Dynamo’s KV-aware router prioritizes finding a worker that already has part of the prompt in its cache. Sessions also remain on the same decoder between turns when this promotes KV reuse.
Mooncake adds a second level of cache in the host’s DRAM. Prefixes evicted from GPU memory can thus be retrieved without being entirely recomputed.
On GLM-5.3 running on GB200 NVL72, Prime targets 100 end-to-end tokens per second per user. In its tested configuration, a ratio of one prefill group to four decode groups achieves 101 tokens/s/user with 66 concurrent sessions per prefill group. Compressing the KV Cache to Four Bits
The amount of context that can be retained in memory quickly becomes a constraint. Prime therefore compresses the latent MLA portion of the KV cache using NVFP4, which is four bits per value accompanied by an FP8 scale per group of 16 values.
Each cache line is thus reduced from 576 to 352 bytes. For the same memory budget, the total capacity measured by Prime increases by about 50%, from 1.09 to 1.63 million cached tokens per decoder.
A specific sparse-MLA kernel then decompresses the data directly during attention computation, without going through an intermediate FP8 buffer in GPU memory. Prime reports maintaining comparable speed per user with prefix caching enabled. Reducing Transfers Between GPUs
However, the separation between prefill and decode creates another problem: efficiently moving the KV cache between workers.
On an initial NVLink configuration, a request of 200,000 tokens could trigger 32,000 copies on a single TP4 rank. Prime adopted the BLHNC memory layout developed by the vLLM community, which groups more contiguous data before transfer.
In a TP8 test, the number of descriptors dropped from 19,559 to approximately 1,940, while the average transfer time decreased from 146 to 78 ms. These measurements correspond to the configuration tested by Prime and do not represent guaranteed performance for all workloads. Tool Calls Are Controlled During Decoding
Optimizations also target agents that manipulate tools. Prime had observed calls to non-existent functions, missing arguments, or incorrect types when using GLM-5.3 at scale.
A contribution to NVIDIA Dynamo now converts tool definitions into structural constraints. vLLM and xgrammar block tokens during decoding that would produce a name or arguments incompatible with the expected schema.
Today, Prime Inference combines NVIDIA Dynamo, vLLM, Mooncake, and FlashInfer, with contributions upstreamed back to the open-source projects. Announced next steps include batch and asynchronous inference, as well as dedicated deployments of fine-tuned models on reserved capacity.