Models
Speech and text models you can use with Freehand. Some use Generic model behavior; others have a dedicated profile for model-specific controls.
Look for Install in Freehand to find models available through local setup. Other entries describe models you can connect to through your own service, not downloads Freehand manages. Compare local choices →
Check your release. Some profiles listed here may not yet be in your installed version. Support also depends on the backend, server version, and checkpoint. Each entry identifies which model behavior to select in Freehand.
Transcription
Generic model behavior
- Recognition
- Language, context, and temperature where supported
- Vocabulary
- Speaches hotwords; context hints on whisper.cpp
- Audio files
- Completed results; optional streaming on Speaches
Choose Generic model behavior with the appropriate backend profile. Available checkpoints and languages depend on your server. whisper.cpp uses its server-loaded model and does not support file streaming or live dictation in Freehand.
Set up Whisper →
Nemotron 3.5 ASR
NVIDIA · Streaming 0.6B
Dedicated model profile
- Language
- Automatic or 32 locales
- Vocabulary
- Shared terms and strength
- Live mode
- Results and optional captions
Language and vocabulary work in completed and live modes. Audio files keep their own settings.
Set up Nemotron 3.5 ASR →Dedicated model profile
- Completed audio
- Language, context, temperature
- Vocabulary
- Shared terms in completed context
- Live mode
- Automatic language and captions
The realtime setup uses the 1.7B checkpoint with vLLM's realtime architecture. Context and vocabulary apply to completed audio only.
Connect your own service
Set up Qwen3-ASR →Dedicated model profile
- Recognition
- Automatic language and punctuation
- Voice
- Recordings and checkpoints
- Audio files
- Completed results
This profile covers the TDT v3 checkpoint on NeMo-Speech.cpp.
Set up Parakeet TDT v3 →Dedicated model profile
- Language
- 14 languages; English default
- Recognition
- Temperature override
- Audio files
- Optional streamed results
Cohere Transcribe on vLLM does not support context or vocabulary hints.
Connect your own service
Set up Cohere Transcribe →Dedicated model profile
- Live mode
- Results and optional captions
- Language
- Automatic detection
- Completed audio
- Recordings and files
Uses the Voxtral Mini 4B Realtime checkpoint. Audio files retain their own selection.
Connect your own service
Set up Voxtral Mini Realtime →Transcript cleanup
Generic model behavior
- Instructions
- Custom cleanup instructions
- Output
- Generation controls provided by the backend
- Recovery
- Raw transcript on cleanup failure
Use an instruction-following model behind a compatible chat API. Check that your server and model support the options you choose. Generic does not apply S1-mini’s trained style controls.
Connect your own service
Set up Compatible chat models →Dedicated model profile
- Styling
- Casual through formal
- Structure
- Prose or lists
- Context
- General or email
English cleanup after transcription. The raw transcript remains the fallback if cleanup fails.
Set up S1-mini →Text to speech
Generic model behavior
- Voice
- Browse available voices or enter an ID
- Speaking speed
- Adjust speed before generation
- Audio
- Listen or explicitly save the generated WAV
Choose Generic model behavior with Speaches or Kokoro-FastAPI. Voice availability depends on the server and model. These profiles do not offer language overrides, voice-style instructions, or cloning.
Connect your own service
Set up Kokoro →Dedicated model profile
- Voice
- Five preset speakers
- Language
- Up to nine, depending on the server
- Local runtime
- Shares NeMo with transcription
Uses the v2602 checkpoint at normal speed. Japanese and Chinese require the corresponding server language frontends.
Set up MagpieTTS Multilingual 357M →'%20fill-rule='nonzero'%3e%3c/path%3e%3cdefs%3e%3clinearGradient%20id='lobe-icons-qwen-_R_0_'%20x1='0%25'%20x2='100%25'%20y1='0%25'%20y2='0%25'%3e%3cstop%20offset='0%25'%20stop-color='%236336E7'%20stop-opacity='.84'%3e%3c/stop%3e%3cstop%20offset='100%25'%20stop-color='%236F69F7'%20stop-opacity='.84'%3e%3c/stop%3e%3c/linearGradient%3e%3c/defs%3e%3c/svg%3e)
Qwen3-TTS
Qwen · 1.7B CustomVoice
Dedicated model profile
- Voice
- Nine preset speakers
- Language
- Automatic or ten languages
- Voice style
- Natural-language instructions
Requires the 1.7B CustomVoice checkpoint. Voice cloning and VoiceDesign are not supported by this profile.
Connect your own service
Set up Qwen3-TTS →Choosing a model
The backend profile selects the server API. Model behavior selects the controls Freehand applies. A different model size or quantization does not automatically require a different profile.
Model selection guide →Speech model families →