Skip to content

Connect a speech server

Connect a service you already run or a hosted provider using its base URL, model ID, and authentication details. For Freehand to install and manage a service on this computer, follow Local runtimes instead.

Have these details ready before opening the connection editor:

Detail What to enter
Base URL The API prefix, usually https://speech.example.com/v1; not the full operation URL
Backend profile The server software, such as Speaches or vLLM
Model ID The exact ID expected by the server; whisper.cpp uses its already-loaded model
Authentication None or API key, according to the server
Gateway requirements Any additional non-secret transcription headers or custom health path
  1. Open the task’s Settings cog, then Transcription, Cleanup, or Speech in the right options sidebar.
  2. Choose Add connection… in its connection picker. Enter a name, backend profile, base URL, and any required authentication.
  3. Choose Save and return. The saved connection is selected for this task.
  4. Choose the model and model profile, review task options, and choose Save in the options sidebar.
Task Required setup Independent of
Voice transcription Speech-to-text connection, model, microphone, and Voice readiness check Recording shortcut, which is optional
Audio file Its own transcription connection and model Microphone and Voice readiness
Text to speech Speech connection, model, voice, and Enable text to speech Microphone and transcription
Cleanup Optional chat connection, model, and cleanup controls The transcription server

Leave Cleanup off to keep the speech model’s original text. You can also add a server directly in Connections; see Saved connections for reuse, switching, and editing shared server details.

Backend describes the server’s API. Model profile, chosen separately in task settings, describes model behavior and exposes supported controls. Neither choice is inferred from a model name or URL.

Use Backend profiles
Speech to text Generic OpenAI-compatible, Speaches, whisper.cpp, vLLM, NeMo-Speech.cpp
Cleanup Generic OpenAI-compatible, llama.cpp, vLLM
Text to speech Generic OpenAI-compatible, Speaches, Kokoro-FastAPI, vLLM-Omni, NeMo-Speech.cpp

Choose the profile matching the server. Use Generic OpenAI-compatible for another compatible service, then check its requirements. Server version and model determine actual support; enabling a use does not verify it.

Location Consider when…
This computer You want inference to use this computer’s resources
Another computer or private server You want to preserve local memory/GPU capacity or share a server
Hosted provider You want a compatible service without managing its runtime

A local GPU is optional. Network speed and server capacity affect response time. Audio or text goes to the endpoint you choose, whose data policy applies. Freehand is free forever; provider usage and server hosting costs are separate.

Tasks can share a gateway or use different servers. Each task keeps its own model and options; tasks sharing a connection share its address and authentication.

Save one connection and enable the operations the gateway exposes. Select it independently for transcription, cleanup, or speech.

One saved gateway, independent task models
Transcription ─┐
Cleanup ───────┼─ https://speech.example.com/v1
Speech ────────┘

A metadata check does not verify every enabled operation.

For http://127.0.0.1:8000/v1 or a private address such as http://192.168.1.50:8000/v1, explicitly enable Allow HTTP for this connection. Freehand otherwise rejects plaintext requests.

Choose Check connection using saved settings. The test reads a health route or model-list metadata; it never sends audio or a prompt or runs listed models.

Result Meaning
Model list received Valid model inventory returned; review whether your selected ID is listed
Server reachable The configured health endpoint responded successfully
Invalid response The response was incompatible, including an HTML gateway page with status 200
Selected model Not listed Your ID was absent, including from an otherwise valid empty inventory

Checks show HTTP status, latency, and a failure category. They cannot prove a model can transcribe or synthesize, or that an endpoint enforces the supplied key. See connection-check results.

Use a custom health path

The optional path is appended beneath the base URL path. A base URL of https://speech.example.com/v1 plus /health checks https://speech.example.com/v1/health, not the host-root route. Include the leading slash. A failed custom check is reported without falling back to model listing. Leave it blank for the backend default: /health for whisper.cpp or model-list metadata for other backends.

Initial Voice setup requires Check connection before Finish setup. Afterward, the saved Voice connection is checked on launch and when relevant connection settings change. Correct a failed check and retry.

Open the task’s Transcription → Spoken language and search by language or code. Server default, Automatic detection, named languages, and custom values depend on the selected model. This requests transcription in the source language, not translation.

Save options before starting a new request. Voice shows context with model options and temperature under Request settings; Audio file groups them in Transcription controls. Only supported controls are offered.

Control Use
Transcription context Short subject matter or expected wording; availability varies by backend, model, and mode
Vocabulary Shared names and phrases, enabled separately for Voice and files; see Vocabulary
Override temperature A value from 0 to 1 where supported; leave off for the server default

Context and vocabulary guide recognition rather than editing its result. Context is limited to 8,192 UTF-8 bytes, including vocabulary when sent as context; your model may accept less. These hints are saved locally and sent to the selected server when used. Keep secrets out of them.

To compare an option, explicitly transcribe a short sample with your selected model, change one setting, and repeat. A rejected request is not retried with different options automatically. Use Cleanup to edit text.

The Speaches launch guide covers GPU/CPU Docker setup, installing one chosen model, and connecting Freehand. For other deployments, follow the runtime’s own guide.