Miso One, the expressive text-to-speech from Miso Labs
Miso Labs launches Miso One, an 8-billion-parameter text-to-speech model with 110ms latency, emotional expression, and voice cloning on Hugging Face.
One hundred ten milliseconds of latency for a voice that laughs, hesitates, or becomes sad: Miso One, the first public model from Y Combinator-backed startup Miso Labs, approaches speech synthesis through emotion.
Boasting 8 billion parameters, the model generates speech with realistic conversational timing and reacts to audio context by aligning with the tone of its interlocutor, from a whisper to a scream, a profile designed for real-time voice agents as well as for narration or podcasting. A few seconds of audio accompanied by their transcription are enough to clone a voice or extend it in the same style, with a demonstration featuring a synthetic version of Sal Khan explaining mathematics.
Under the hood, the architecture combines Sesame CSM and a Residual Vector Quantization with 32 codebooks, a 7.7 billion parameter backbone inspired by Llama 3.2 being coupled with a 300 million parameter decoder. The weights are published on Hugging Face and the inference code on GitHub, allowing the model to be run locally on a sufficiently equipped GPU, with an online demo at misolabs.ai and a paid API completing the upcoming offering. Founders Aoden Teo and Cassidy Dalva aim, in their words, for the world's most emotive voice foundation models.