Cloud transcription, local cleanup: Superwhisper structures its S1 model family
Superwhisper now divides dictation processing among S1-Voice for speech recognition, S1-Language for advanced transformations, and S1-mini for local cleanup.
Superwhisper has introduced three internally developed models covering different stages of its dictation tool. S1-Voice converts speech into text in the cloud, S1-Language transforms that text based on instructions, and S1-mini handles more focused cleanup and formatting tasks locally.
This distinction matters: S1-mini is not a speech recognition model. It receives an existing transcript and then removes hesitations, interprets spoken corrections, structures lists, or formats messages. A fully local setup therefore requires pairing it with an on-device speech model. Superwhisper recommends Cohere Transcribe locally with S1-mini, while its suggested cloud configuration combines S1-Voice with S1-Language.
S1-mini has 600 million parameters and a reported file size of 462 MB. It runs without making network requests. A slider offers five tone settings, ranging from casual lowercase writing to formal text with full punctuation and expanded contractions.
The model can turn a dictated sequence into a list, recognize the structure of an email, convert a spoken address into a properly formatted email address, or process a correction such as “Tuesday, I mean Thursday.” Superwhisper says it will not add information that was never spoken, verify facts, soften profanity, flag subject matter, or alter a dialect. Its role is limited to preserving and organizing dictated content.
The company reports results from an internal evaluation covering 7,519 cases drawn from 104 transcripts withheld from training. It claims 94.8% token-level accuracy, an 11.6% text-edit error rate, and correct selection of the expected structure—list or paragraph—in 97.6% of cases. Exact email-address reproduction reached 92%. These figures primarily measure the post-processing of expected text, not the quality of the speech recognition stage that precedes it.
Superwhisper also says it is releasing S1-mini with open weights on Hugging Face. However, the announcement does not provide a direct repository link or specify the license. The company’s public model directory did not list S1-mini at the time of publication. The permitted uses for modification, redistribution, and commercial deployment therefore remain unclear.
S1-Voice handles the first stage of the pipeline. Hosted by Superwhisper, it is available to Pro subscribers and supports more than 100 languages, according to the company’s voice model documentation. Superwhisper says most dictations under 30 seconds are returned 0.32 seconds after the user stops speaking.
Across eight datasets covering meetings, earnings calls, and spontaneous speech, the company reports an average word error rate of 6.8% and a score of 2.2% on LibriSpeech. S1-Voice reportedly ranked first among the 15 models tested, with a blended score of 83 out of 100, compared with 76 for Wispr Flow. These evaluations were conducted by Superwhisper. The announcement does not provide the complete audio files, outputs, settings, or protocol required for independent reproduction, and its blended score is not a standard metric that can be directly compared with other rankings.
S1-Language operates after transcription to apply more advanced instructions. It can clean up dictation, summarize a meeting, follow a predefined note format, or adapt an email to specific rules. This model is also cloud-hosted, restricted to the Pro tier, and listed with a 128,000-token context window. It appears in Superwhisper’s model picker alongside systems from Anthropic, OpenAI, and Groq.
Model selection determines how data is processed. With a fully local pipeline, both audio and text remain on the device. When a cloud model is selected, data passes through Superwhisper’s infrastructure. The company states that content sent to its models or integrated providers is neither retained nor used for training under zero-data-retention agreements. These are contractual commitments described by the company; they do not turn a cloud configuration into local processing.
The August 19 announcement formalizes a family that had already been rolling out gradually. The company’s changelog indicates that S1-Voice and S1-Language left experimental status in April 2026. S1-mini became available experimentally in June, with tone controls added in July. Version 2.17.2, released on August 6, then officially listed the arrival of the S1 family.
Superwhisper is available on macOS, Windows, and iOS, but its public documentation does not yet clearly specify S1-mini compatibility across every platform and processor type. Beyond the company’s internal results, further evaluation will be needed across accents, noisy environments, multilingual dictation, and the specialized terminology these models are intended to handle.