Skip to content

Backend compatibility

Choose a backend profile for the API your server exposes. Choose a model profile separately for the languages and controls of its loaded model. The server can run on your computer, a LAN server, a VPS, or a hosted service.

Use the recipes below when you manage the service yourself, need other models or runtime settings, or run the server on another machine.

Server Launch instructions Freehand operation
whisper.cpp Native Windows or Docker Transcription with a server-loaded model
vLLM Windows + Docker, NVIDIA GPU Qwen3-ASR transcription or S1-mini cleanup
Speaches Docker GPU or CPU Transcription; optional speech playback
Kokoro-FastAPI Docker CPU or GPU Speech playback with voice discovery
llama.cpp Native Windows server S1-mini or compatible custom cleanup
NeMo-Speech.cpp Runtime installation Completed/live transcription and MagpieTTS speech
vLLM-Omni Qwen3-TTS serving guide Text to speech with preset voices
Generic OpenAI-compatible Connect an existing endpoint The operations implemented by your server
  • Shell and containers: examples use PowerShell and loopback ports. Docker recipes require Linux containers; NVIDIA examples need working WSL2 GPU access. Follow Docker’s GPU setup if docker run --gpus ... cannot access the card.
  • Downloads and storage: first use downloads the image and selected model. Stopping a container keeps weights in its recipe’s folder or named volume.
  • Resources and shutdown: leave memory for every concurrent model or stop unused servers. Recipes include stop/restart commands or link upstream setup.

For a remote or hosted deployment, use the same profile with the administrator’s HTTPS base URL, model ID, and authentication. Loopback recipes serve a client on the same computer. See connection topologies for separate transcription, cleanup, and speech endpoints.

Action What it checks
Check connection Health or model metadata
Refresh voices Voice metadata on supported speech servers
Short recording or voice preview Runs only the model you explicitly selected so you can review its output

Health and discovery never generate audio or run transcription.

Where Configure
Connections on the activity rail Endpoint, authentication, backend profile, and allowed uses
Workflow Settings cog Active connection, model, model profile, voices, and model options
Local runtimes Installation and loaded models for managed connections

Voice transcription, Audio-file transcription, Cleanup, and Text to speech select their own connection. Tasks using one managed runtime share its selected model and get the matching profile automatically. Manual model options are remembered per connection and model.

Transcription context and temperature depend on the selected backend and model profile. Vocabulary supplies recognition terms through the supported request fields. See the language guide for language selection.

Voice-style instructions are available with Qwen3-TTS on vLLM-Omni. Translation, timestamps, diarization, reference-audio inputs, and progressive speech playback are not available. A server may offer these features without Freehand exposing them.

Use the protocol reference for exact request fields and safety limits.