Pocket TTS now generates each voice fragment in a single pass with Drifting

Kyutai adapts the Drifting generative objective to Pocket TTS, its 100-million-parameter speech synthesis model. This new method maintains single-pass generation, achieves a 0.96% WER, and avoids the Jacobian computations used by its previous LSD training.

A New Training Method for Pocket TTS

Every 80 ms, Pocket TTS must produce the next latent chunk that will compose the voice. Rather than multiplying generation steps as in a traditional diffusion process, its sampler head directly transforms noise into the next latent in a single pass.

Kyutai previously used Lagrangian self-distillation, or LSD, to train this part of the model. The laboratory is now experimenting with Drifting, a recent generative objective that shifts iteration to the training phase rather than inference.

According to Kyutai, this is the first known application of Drifting to a speech model and an autoregressive model. Attracting Generations Toward the Data

The principle differs from classic flow matching. During training, multiple candidates are generated for the next audio chunk. Drifting pulls them closer to the real data while pushing them away from each other when they become too similar.

Once the generated distribution aligns with that of the data, these two forces balance out.

This construction avoids the Jacobian-vector products required with LSD. Kyutai describes them as expensive and numerically delicate during training. No iterative solver or second network is added to the generation.

Pocket TTS therefore retains its single-pass operation for each audio chunk. A Learned Rather Than Fixed Temperature

However, the adaptation did not work out of the box with the original Drifting recipe. The main change concerns the kernel temperature, which determines the distance at which samples interact during training.

Fixing it from the start at its final value caused the candidates to collapse and kept the WER between 15% and 20%. Kyutai therefore made this temperature learnable. In its various successful trials, it converges to a value between 0.055 and 0.057.

Two other parameters complete the recipe: distance normalization for each frame and sufficiently large batches. With 64 rows per step, the model did not reach the targeted quality level within the tested window. The laboratory settled on 128 rows and 32 to 64 negative candidates. 0.96% WER Compared to 0.91% for LSD

On 1,127 items from LibriSpeech test-clean, using the same backbone architecture and codec, the Drifting version achieves a 0.96% WER, an UTMOS score of 4.32, and a speaker similarity of 0.926.

The LSD baseline reaches 0.91%, 4.31, and 0.927 respectively. Kyutai considers these differences to be within the evaluation noise. Both variants also maintain the same inference cost.

Another evaluation of the `englishdrifting26-09` model, with different parameters, measures a 0.90% WER and 4.37 UTMOS, compared to 0.90% and 4.36 for the default English model. The Cost Shifts to Training

While Drifting simplifies certain calculations, it does not yet reduce the overall training cost. On four H100 GPUs, the best Drifting recipe reaches the 4.0 UTMOS threshold after about 10 to 12 hours, compared to 4.9 to 5.6 hours with LSD.

At comparable quality, Kyutai currently estimates the training cost to be roughly twice that of LSD.

However, the model remains intended for local execution on CPU. Pocket TTS has 100 million parameters, runs faster than real-time without a GPU, and supports voice cloning from an audio clip. The experimental variant is available under the name `englishdrifting26-09`, with its code and training recipe published in the repository.