Skip to content

Get started

Freehand turns speech into text and text into speech using a service you choose. Install the app, connect a service or set up a local runtime, then try one task. You can use each task independently.

  • A supported computer: Windows 11 x64 or ARM64 with WebView2, or macOS 13+ on Apple Silicon or Intel.
  • Somewhere to run speech: a local runtime managed by Freehand, or a compatible service on your computer, another machine, or a hosted provider.
  • For dictation: a microphone. Audio files and text to speech do not need one.

Freehand is free and open source, with no account or subscription. Your chosen provider or hosting may charge for use. A local GPU is not required when you connect to another machine or a hosted service. On Windows ARM64, choose a compatible service; managed Windows runtimes currently require x64. See Windows setup for architecture choices.

Freehand opens Voice with Set up voice transcription. Choose Audio file or Text to speech on the left if you want to start with a different task.

Choose where processing will run:

Freehand can install and manage a supported runtime on this computer. Runtime binaries and models are separate, explicit downloads; neither is bundled with the app.

  1. Open Local runtime. Follow local runtime setup to install NeMo and download the recommended Nemotron 3.5 Streaming model.
  2. Start the runtime. Wait for Running, then select its built-in NeMo-Speech.cpp Connection in Voice. No URL or API key is needed.
  3. Enable realtime for live dictation. Turn on Realtime transcription in Voice’s Transcription options and finish microphone setup below.

For speech generation, add MagpieTTS to NeMo. The local runtime guide also covers whisper.cpp transcription and optional S1-mini cleanup.

Each task selects its own connection. To share a manual service, enable the needed connection uses and select it in each task. A managed connection supplies its runtime’s selected model.

For a manual connection, enter the base URL, not the full transcription route. Most transcription backends use a /v1 prefix; whisper.cpp uses the server root. Follow the URL examples for your backend. HTTPS is the default. Allow HTTP only on a trusted local or LAN connection: it sends audio and credentials without encryption.

Open your task’s Settings cog, then Transcription, and choose a supported language, Server default, or Automatic detection. Language selection explains which choice to use. S1-mini cleanup supports English only.

Leave Cleanup off for your first attempt to receive the speech model’s transcript unchanged. You can add cleanup later; if it fails, Freehand falls back to the raw transcript.

  1. Finish Voice setup. Select your connection and model, then follow Next step to check the microphone and shortcut. On macOS, grant the requested permissions.
  2. Check the connection. Choose Check connection, resolve any reported problem, then Finish setup and save. The check reads metadata; it does not submit audio or run a model.
  3. Focus the destination. Click the text field in the application where your words should appear.
  4. Record a short phrase. Press Ctrl + Shift + Space to start, then press it again to stop. If you changed the shortcut, use that chord.
  5. Keep the destination focused. Wait for transcription and delivery.

Expected result: your text appears in the destination. If insertion is blocked, check for partially inserted text, choose Copy, and paste it yourself. Freehand will not redirect the result to a different focused window. See safe text insertion.

Open Audio file, select its transcription connection and model, then Choose a file. Selecting a file does not upload it; Transcribe starts the request. Copy the finished result yourself.

Open Text to speech → Configure speech, select a speech connection, model, and voice, then turn on Enable text to speech and Save. Enter text and choose Speak. Playback is optional and off by default.

Start with one successful request, then adjust the options you need.

Feature Initial setting Learn more
Local voice detection and status overlay On Voice controls
Silence trimming, automatic stop, pause-aware checkpoints Off Recording options
Transcript cleanup Off Cleanup
Memory-only transcript history Off History
Audio-file text updates Completed output; streaming optional Audio files
Text-to-speech playback Off Text to speech
Insert dictation into the original application On, when focused and safe Delivery rules
Automatic update checks On Workspace and settings
Windows Mica material Off Workspace appearance

Voice detection alone does not enable automatic stop, trimming, or checkpoints.