Fish Audio raises $52M one year after leaving GitHub
Fish Audio raises $52 million to establish its S2.1 Pro voice model. A rapid cloning technology targeting the enterprise market.
Fish Audio closes a fifty-two million dollar seed round, one year after its incorporation. Coreline Ventures and Capital Today lead the round, joined by 359 Capital, Play Time, HF0, 645 Ventures, Parable, Bayhouse Ventures, Carya Venture Partners, and Alphalist Partners. The origin of the project stems from a rather specific image. Shijia Liao, then a video research engineer at NVIDIA and an avid viewer of VTuber streams, could hardly tolerate the expressiveness-lacking synthetic voices and began training models on a single 4090 card at home. The resulting repository, Fish Speech, now exceeds thirty-one thousand stars on GitHub.
The company's thesis lies in a distinction. Pronouncing correctly and performing accurately are two different problems, and the industry has allegedly only tackled the former. A synthetic voice convinces on one sentence and cracks by the tenth. Its co-founder and CEO, who previously worked on voice AI at Amazon Alexa and then Meta, claims to have approached the problem from the opposite end.
Five models were released in one year, four for synthesis and one for transcription, with the headcount growing from three to twenty-two people. One point deserves clarification, which the company's communication leaves in the dark: three of the synthesis models are open, but the latest, S2.1 Pro, is only accessible via the paid API. The latter covers more than eighty-three languages with native prosody and word-level emotion control, via more than fifteen thousand natural language commands, and clones a voice from five seconds of audio.
Post-training relies on actual user preferences, which reportedly explains the model's performance on accented English, Mandarin, Japanese, Korean, and Spanish, where models tuned to standard American English fall short. The research team has also overhauled the inference stack with in-house FP8 kernels, claiming over eight thousand tokens per second on a single H200. On-premise deployment, zero data retention, and HIPAA-compliant configurations target regulated sectors, with enterprise and developers accounting for two-thirds of an announced twenty-one million dollars in annual recurring revenue.
The next steps will focus on an audio understanding model and speech-to-speech. Nonetheless, the highlighted comparisons, with about two-thirds of listeners preferring S2.1 Pro to its competitors in blind listening tests, originate from the company and are not accompanied by any published protocol.