Live transcription
Live transcription streams microphone audio to your selected server while you speak. Review the preview in Freehand or enable overlay captions. Stop to finalize the text, run optional cleanup, and deliver it while the original destination is still safe. Manual-copy mode leaves it ready to copy.
Choose a model and backend
Section titled “Choose a model and backend”| Model profile | Qualified backend | Live recognition controls |
|---|---|---|
| Nemotron 3.5 ASR streaming | NeMo-Speech.cpp v0.1.0 | Automatic/explicit language; shared vocabulary and strength |
| Qwen3-ASR | vLLM v0.28.0 | Automatic language detection |
| Voxtral Mini Realtime | vLLM v0.28.0 | Automatic language detection |
For local setup on Windows or macOS, install managed NeMo with Nemotron. For a server you manage elsewhere, follow the chosen model’s guide. Select its explicit model profile; a compatible URL or model name alone does not enable live mode.
Enable live mode
Section titled “Enable live mode”- Open Voice transcription → Settings → Transcription in the right sidebar.
- Select the connection, loaded model, and model profile.
- Enable Realtime transcription, shown for a supported combination.
- Optionally enable Live overlay captions. The main Overlay preference must also be on.
- Choose Save, then focus the destination and start with your recording shortcut.
Language and vocabulary controls follow the model profile. Unavailable completed controls remain saved. Edit shared terms under Settings → Vocabulary or Voice’s Vocabulary options.
While recording
Section titled “While recording”| What you do | What happens |
|---|---|
| Speak | The preview can change; overlay captions show the newest words |
| Stop or release hold-to-talk | Final text replaces the preview and can proceed to cleanup and delivery |
| Cancel | The preview is discarded |
| Lose the stream | The preview is discarded; start a new recording after fixing the connection |
| Turn realtime off for the next run | The same connection/model remain, with saved completed capture preferences |
Live mode uses your toggle/hold shortcut and duration limit. Silence trimming, checkpoints, and automatic stop are bypassed. Keep the destination focused until final delivery; safe insertion rules still apply.
Audio files keep independent connection, language, and options. When Voice and files use the same managed runtime they share its selected transcription model; otherwise their models are independent. File requests remain completed uploads, even when the server streams text back.