Fish Audio brings its voice studio to iPhone and Android
Fish Audio launches its voice studio on iPhone and Android, with 2 million voices, cloning, and dialogue in more than 80 languages.
Fish Audio is moving beyond the browser with a mobile app dedicated to voice creation. First launched on iPhone, it provides access to the voice library, dialogue generation, and voice creation tools from a phone. An Android version followed shortly afterward with a similar set of features.
The app lets users write or paste a script, assign a different voice to each section, and adjust emotion, speed, and volume. Results can be played, saved, shared, and reopened from the project history. Fish Audio is targeting voiceovers for videos, podcasts, audiobooks, and stories with multiple characters.
Users can browse a community library advertised as containing more than two million voices. That figure refers to the catalog available on the platform, not to two million voices created or individually verified by Fish Audio. Usage rights therefore depend on each voice, whether it is public or private, and the permissions granted by its creator.
The mobile studio uses features from S2.1 Pro, the company’s current recommended speech-generation model. It supports multi-speaker dialogue and 83 automatically detected languages. Written directions placed in brackets can guide the delivery at a specific point in the text, such as requesting a whisper, a laugh, a hesitation, or a nervous tone.
These directions are not restricted to a fixed list. Fish Audio treats them as natural-language descriptions, allowing for more precise instructions than selecting a basic emotion from a menu. The app brings this approach into a touch-based interface without requiring users to work with the API.
There are two ways to create a new voice. The first involves recording or importing an authorized sample to make a voice clone. The second, Voice Design, generates a voice from a description of characteristics such as apparent age, accent, vocal tone, or delivery style. The resulting voice can then read a text or be assigned to one of the speakers in a dialogue.
The phrase “full voice studio” used in the announcement requires some qualification. The current store listings mainly describe speech generation, dialogue creation, cloning, Voice Design, and project management. They do not yet confirm the inclusion of every product available on the website, such as transcription, audio translation, track separation, sound effects, or voice changing.
The official App Store version is published by Hanabi AI Inc., the company behind Fish Audio. This distinction matters because several third-party apps already use very similar names and descriptions. The official app is 82.8 MB, requires iOS 16.4 or later, and is designed for iPhone. It can run on some Macs equipped with Apple silicon, but Apple does not list it as a verified macOS app.
The 83 supported languages apply to voice generation, not the interface. At launch, the App Store lists only English, Japanese, Simplified Chinese, and Traditional Chinese as interface languages. Users can therefore generate speech in French without having access to a French-language interface.
The app is free to download, but generation remains subject to Fish Audio’s credit system. The service’s current free plan includes 8,000 monthly credits, representing up to roughly seven minutes of generation, with a limit of 500 characters per request. Subscriptions and credit packs are sold through the app. Prices charged by Apple may differ from promotions displayed on the website.
Moving to mobile does not mean the model runs directly on an iPhone or Android device. Fish Audio does not mention an offline mode or on-device processing, so the app acts as an interface to its remote service. According to the information provided to Apple, audio recordings, other uploaded content, search history, purchases, and several identifiers may be linked to the user’s account.
This first release therefore adds a new way to access the service rather than introducing a new voice model. Its main benefit is the ability to record a voice, build a dialogue, and share the result from the same device without returning to the web studio.