Backend compatibility
Choose a backend profile for the API your server exposes. Choose a model profile separately for the languages and controls of its loaded model. The server can run on your computer, a LAN server, a VPS, or a hosted service.
Run a backend
Section titled “Run a backend”Use the recipes below when you manage the service yourself, need other models or runtime settings, or run the server on another machine.
| Server | Launch instructions | Freehand operation |
|---|---|---|
| whisper.cpp | Native Windows or Docker | Transcription with a server-loaded model |
| vLLM | Windows + Docker, NVIDIA GPU | Qwen3-ASR transcription or S1-mini cleanup |
| Speaches | Docker GPU or CPU | Transcription; optional speech playback |
| Kokoro-FastAPI | Docker CPU or GPU | Speech playback with voice discovery |
| llama.cpp | Native Windows server | S1-mini or compatible custom cleanup |
| NeMo-Speech.cpp | Runtime installation | Completed/live transcription and MagpieTTS speech |
| vLLM-Omni | Qwen3-TTS serving guide | Text to speech with preset voices |
| Generic OpenAI-compatible | Connect an existing endpoint | The operations implemented by your server |
Before running a local recipe
Section titled “Before running a local recipe”- Shell and containers: examples use PowerShell and loopback ports. Docker
recipes require Linux containers; NVIDIA examples need working WSL2 GPU
access. Follow Docker’s GPU setup
if
docker run --gpus ...cannot access the card. - Downloads and storage: first use downloads the image and selected model. Stopping a container keeps weights in its recipe’s folder or named volume.
- Resources and shutdown: leave memory for every concurrent model or stop unused servers. Recipes include stop/restart commands or link upstream setup.
For a remote or hosted deployment, use the same profile with the administrator’s HTTPS base URL, model ID, and authentication. Loopback recipes serve a client on the same computer. See connection topologies for separate transcription, cleanup, and speech endpoints.
Check your connection
Section titled “Check your connection”| Action | What it checks |
|---|---|
| Check connection | Health or model metadata |
| Refresh voices | Voice metadata on supported speech servers |
| Short recording or voice preview | Runs only the model you explicitly selected so you can review its output |
Health and discovery never generate audio or run transcription.
Choose profiles per operation
Section titled “Choose profiles per operation”| Where | Configure |
|---|---|
| Connections on the activity rail | Endpoint, authentication, backend profile, and allowed uses |
| Workflow Settings cog | Active connection, model, model profile, voices, and model options |
| Local runtimes | Installation and loaded models for managed connections |
Voice transcription, Audio-file transcription, Cleanup, and Text to speech select their own connection. Tasks using one managed runtime share its selected model and get the matching profile automatically. Manual model options are remembered per connection and model.
Beyond the current contract
Section titled “Beyond the current contract”Transcription context and temperature depend on the selected backend and model profile. Vocabulary supplies recognition terms through the supported request fields. See the language guide for language selection.
Voice-style instructions are available with Qwen3-TTS on vLLM-Omni. Translation, timestamps, diarization, reference-audio inputs, and progressive speech playback are not available. A server may offer these features without Freehand exposing them.
Use the protocol reference for exact request fields and safety limits.