Skip to content

Speech model families

A model family groups related recognition or speech-generation models. Choose both a model and a backend that can serve it, then select the matching Freehand profile from the table below. A different size, quantization, or server alias does not automatically change the profile; check the model guide’s requirements.

Family Backend Freehand model profile
Whisper, Distil-Whisper Speaches, whisper.cpp, compatible transcription servers Generic
Nemotron 3.5 streaming NeMo-Speech.cpp Nemotron
Parakeet TDT v3 NeMo-Speech.cpp Parakeet TDT v3
Qwen3-ASR vLLM Qwen3-ASR
Cohere Transcribe vLLM Cohere Transcribe
Voxtral Mini 4B Realtime vLLM Voxtral Mini Realtime
Kokoro Speaches, Kokoro-FastAPI Generic
MagpieTTS Multilingual 357M v2602 NeMo-Speech.cpp MagpieTTS
Qwen3-TTS 1.7B CustomVoice vLLM-Omni Qwen3-TTS
S1-mini llama.cpp, vLLM, compatible chat servers S1-mini
Family, runtime, or model format?

Whisper is a model family; faster-whisper and whisper.cpp are different implementations. GGUF, ONNX, and MLX describe packaging rather than new Freehand behavior profiles. A different size, quantization, or alias still needs a compatible backend and model profile. The server can run on any machine reachable from the Windows or macOS client.

Recognition family Distinct behavior to explore
Microsoft VibeVoice ASR Streaming Streaming recognition with speaker-attributed output and hotword context. Its plugin server uses a different protocol from Freehand’s vLLM adapter.
NVIDIA Canary Multilingual transcription and speech translation.
SenseVoice Language identification, text normalization, and additional audio annotations in the FunASR ecosystem.
Kyutai STT Streaming recognition with dedicated English and English/French variants.
IBM Granite Speech A speech family with distinct checkpoint architectures, including English TurboCTC.
IndicConformer Dedicated recognition across 22 Indian languages.

For speech generation, Chatterbox, Fish Audio S2, and OmniVoice offer different expression, voice, and language controls. XTTS, F5-TTS, CosyVoice, Piper, VoxCPM, and Higgs are also distinct families worth considering when choosing a speech server.

Some servers can already provide ordinary transcription or speech generation through a Generic connection if they accept its request fields and response formats. Generic does not enable a model’s specialized controls or an unrelated streaming API.

Use Hugging Face’s recognition and speech-generation directories to discover families, then check:

  • The task: general transcription, alignment, diarization, and language-specific derivatives can appear together.
  • The serving API: popularity and download counts do not establish compatibility.
  • The controls: choose the model and backend together for the workflow you need.