Models

Speech and text models you can use with Freehand. Some use Generic model behavior; others have a dedicated profile for model-specific controls.

Look for Install in Freehand to find models available through local setup. Other entries describe models you can connect to through your own service, not downloads Freehand manages. Compare local choices →

Check your release. Some profiles listed here may not yet be in your installed version. Support also depends on the backend, server version, and checkpoint. Each entry identifies which model behavior to select in Freehand.

Transcription

Whisper

Whisper & Distil-Whisper · Completed transcription

Generic model behavior

Recognition
Language, context, and temperature where supported
Vocabulary
Speaches hotwords; context hints on whisper.cpp
Audio files
Completed results; optional streaming on Speaches

Choose Generic model behavior with the appropriate backend profile. Available checkpoints and languages depend on your server. whisper.cpp uses its server-loaded model and does not support file streaming or live dictation in Freehand.

Install in Freehand with whisper.cpp →Windows onlyWhisper Tiny through Large, including Turbo and quantized variants
Set up Whisper →

Qwen3-ASR

Qwen · Speech recognition

Dedicated model profile

Completed audio
Language, context, temperature
Vocabulary
Shared terms in completed context
Live mode
Automatic language and captions

The realtime setup uses the 1.7B checkpoint with vLLM's realtime architecture. Context and vocabulary apply to completed audio only.

Connect your own service

Set up Qwen3-ASR →

Cohere Transcribe

Cohere · Speech recognition

Dedicated model profile

Language
14 languages; English default
Recognition
Temperature override
Audio files
Optional streamed results

Cohere Transcribe on vLLM does not support context or vocabulary hints.

Connect your own service

Set up Cohere Transcribe →

Voxtral Mini Realtime

Mistral · Streaming recognition

Dedicated model profile

Live mode
Results and optional captions
Language
Automatic detection
Completed audio
Recordings and files

Uses the Voxtral Mini 4B Realtime checkpoint. Audio files retain their own selection.

Connect your own service

Set up Voxtral Mini Realtime →

Transcript cleanup

Compatible chat models

Instruction-following text models · Optional cleanup

Generic model behavior

Instructions
Custom cleanup instructions
Output
Generation controls provided by the backend
Recovery
Raw transcript on cleanup failure

Use an instruction-following model behind a compatible chat API. Check that your server and model support the options you choose. Generic does not apply S1-mini’s trained style controls.

Connect your own service

Set up Compatible chat models →

Text to speech

Kokoro

Speech generation · Preset voices

Generic model behavior

Voice
Browse available voices or enter an ID
Speaking speed
Adjust speed before generation
Audio
Listen or explicitly save the generated WAV

Choose Generic model behavior with Speaches or Kokoro-FastAPI. Voice availability depends on the server and model. These profiles do not offer language overrides, voice-style instructions, or cloning.

Connect your own service

Set up Kokoro →

Qwen3-TTS

Qwen · 1.7B CustomVoice

Dedicated model profile

Voice
Nine preset speakers
Language
Automatic or ten languages
Voice style
Natural-language instructions

Requires the 1.7B CustomVoice checkpoint. Voice cloning and VoiceDesign are not supported by this profile.

Connect your own service

Set up Qwen3-TTS →

Choosing a model

The backend profile selects the server API. Model behavior selects the controls Freehand applies. A different model size or quantization does not automatically require a different profile.

Model selection guide →
Speech model families →