Skip to content

Transcript post-processing

Turn on Cleanup to edit a transcript before delivery. Voice and audio files share this optional stage. It sends transcript text to your selected chat service, separate from speech recognition; it does not send the original audio.

Desired result Configuration
Leave transcription unchanged Keep Cleanup off; no chat connection is needed
Edit using your instructions Choose a compatible chat connection/model and the Generic model profile
Use English cleanup presets Choose S1-mini by Superwhisper, through managed llama.cpp or a supported manual service
  1. Open Voice transcription or Audio file → Settings → Cleanup. Select a saved cleanup connection, or Add connection… using the server connection guide.
  2. Enable cleanup and choose a model. Refresh models reads metadata; it does not test cleanup quality. You can also enter the exact model ID.
  3. Choose Generic under Model profile. Write the Custom system instruction, or restore the recommended instruction as a starting point.
  4. Save, then try a short transcription. Review the cleaned result before relying on it for longer text.

Write how to edit the text without a transcript placeholder: Freehand sends the raw transcript separately with the selected connection’s credentials. Instructions are saved locally; API keys stay in Windows Credential Manager or macOS Keychain.

For manual connections, choose the model profile explicitly. Freehand does not identify S1-mini from its name. Switching profiles preserves custom instructions and S1-mini choices for when you switch back.

Follow local llama.cpp setup, then select its built-in connection in Cleanup. The runtime guide lists supported platforms. Local transcription and cleanup use separate runtimes and can run together.

S1-mini requires English and reasoning off. Review the language behavior below before using automatic language detection.

Open Cleanup → Generation controls and save before the next recording or file job. Active work retains its settings even if cleanup has not started yet.

Control Values and effect
Limit output tokens Optional for Generic, llama.cpp, and vLLM; starts at 2,048 and accepts 1–65,536 whole tokens, not words
Output limit off Uses the server’s limit; keeps a valid value for later
Disable reasoning with Generic model profile Optional with llama.cpp/vLLM; asks for a result without a thinking phase. Off uses server behavior
S1-mini Required off Freehand requests reasoning off with llama.cpp/vLLM
S1-mini Disable on server With a Generic backend, configure reasoning off yourself

An output limit neither enlarges the context window nor splits long inputs; request timeout is independent. Allow room for the complete cleaned transcript. Support for optional reasoning control depends on the server and model template. Turn a saved custom reasoning override off before switching the backend from llama.cpp/vLLM to Generic.

See llama.cpp generation controls for qualified versions and model limits.

Choose vLLM as the connection’s backend profile for a supported v0.28.0 server. Configure its /v1 endpoint and text model separately from transcription. The same generation controls apply, including required reasoning off for S1-mini. See the vLLM guide for deployment options.

S1-mini input language Behavior
English Eligible for cleanup
Explicit or server-reported non-English / mixed language Skips cleanup and keeps raw text
Unknown Runs assuming English, as displayed in its controls

Custom cleanup language support depends on your model and instruction. Language selection explains the complete S1-mini rules.