Speech model families
A model family groups related recognition or speech-generation models. Choose both a model and a backend that can serve it, then select the matching Freehand profile from the table below. A different size, quantization, or server alias does not automatically change the profile; check the model guide’s requirements.
Families with a Freehand path
Section titled “Families with a Freehand path”| Family | Backend | Freehand model profile |
|---|---|---|
| Whisper, Distil-Whisper | Speaches, whisper.cpp, compatible transcription servers | Generic |
| Nemotron 3.5 streaming | NeMo-Speech.cpp | Nemotron |
| Parakeet TDT v3 | NeMo-Speech.cpp | Parakeet TDT v3 |
| Qwen3-ASR | vLLM | Qwen3-ASR |
| Cohere Transcribe | vLLM | Cohere Transcribe |
| Voxtral Mini 4B Realtime | vLLM | Voxtral Mini Realtime |
| Kokoro | Speaches, Kokoro-FastAPI | Generic |
| MagpieTTS Multilingual 357M v2602 | NeMo-Speech.cpp | MagpieTTS |
| Qwen3-TTS 1.7B CustomVoice | vLLM-Omni | Qwen3-TTS |
| S1-mini | llama.cpp, vLLM, compatible chat servers | S1-mini |
Family, runtime, or model format?
Whisper is a model family; faster-whisper and whisper.cpp are different implementations. GGUF, ONNX, and MLX describe packaging rather than new Freehand behavior profiles. A different size, quantization, or alias still needs a compatible backend and model profile. The server can run on any machine reachable from the Windows or macOS client.
Other families to consider
Section titled “Other families to consider”| Recognition family | Distinct behavior to explore |
|---|---|
| Microsoft VibeVoice ASR Streaming | Streaming recognition with speaker-attributed output and hotword context. Its plugin server uses a different protocol from Freehand’s vLLM adapter. |
| NVIDIA Canary | Multilingual transcription and speech translation. |
| SenseVoice | Language identification, text normalization, and additional audio annotations in the FunASR ecosystem. |
| Kyutai STT | Streaming recognition with dedicated English and English/French variants. |
| IBM Granite Speech | A speech family with distinct checkpoint architectures, including English TurboCTC. |
| IndicConformer | Dedicated recognition across 22 Indian languages. |
For speech generation, Chatterbox, Fish Audio S2, and OmniVoice offer different expression, voice, and language controls. XTTS, F5-TTS, CosyVoice, Piper, VoxCPM, and Higgs are also distinct families worth considering when choosing a speech server.
Some servers can already provide ordinary transcription or speech generation through a Generic connection if they accept its request fields and response formats. Generic does not enable a model’s specialized controls or an unrelated streaming API.
Reading model directories
Section titled “Reading model directories”Use Hugging Face’s recognition and speech-generation directories to discover families, then check:
- The task: general transcription, alignment, diarization, and language-specific derivatives can appear together.
- The serving API: popularity and download counts do not establish compatibility.
- The controls: choose the model and backend together for the workflow you need.