Connect a speech server
Connect a service you already run or a hosted provider using its base URL, model ID, and authentication details. For Freehand to install and manage a service on this computer, follow Local runtimes instead.
Information Freehand needs
Section titled “Information Freehand needs”Have these details ready before opening the connection editor:
| Detail | What to enter |
|---|---|
| Base URL | The API prefix, usually https://speech.example.com/v1; not the full operation URL |
| Backend profile | The server software, such as Speaches or vLLM |
| Model ID | The exact ID expected by the server; whisper.cpp uses its already-loaded model |
| Authentication | None or API key, according to the server |
| Gateway requirements | Any additional non-secret transcription headers or custom health path |
Start from your task
Section titled “Start from your task”- Open the task’s Settings cog, then Transcription, Cleanup, or Speech in the right options sidebar.
- Choose Add connection… in its connection picker. Enter a name, backend profile, base URL, and any required authentication.
- Choose Save and return. The saved connection is selected for this task.
- Choose the model and model profile, review task options, and choose Save in the options sidebar.
| Task | Required setup | Independent of |
|---|---|---|
| Voice transcription | Speech-to-text connection, model, microphone, and Voice readiness check | Recording shortcut, which is optional |
| Audio file | Its own transcription connection and model | Microphone and Voice readiness |
| Text to speech | Speech connection, model, voice, and Enable text to speech | Microphone and transcription |
| Cleanup | Optional chat connection, model, and cleanup controls | The transcription server |
Leave Cleanup off to keep the speech model’s original text. You can also add a server directly in Connections; see Saved connections for reuse, switching, and editing shared server details.
Choose a compatibility profile
Section titled “Choose a compatibility profile”Backend describes the server’s API. Model profile, chosen separately in task settings, describes model behavior and exposes supported controls. Neither choice is inferred from a model name or URL.
| Use | Backend profiles |
|---|---|
| Speech to text | Generic OpenAI-compatible, Speaches, whisper.cpp, vLLM, NeMo-Speech.cpp |
| Cleanup | Generic OpenAI-compatible, llama.cpp, vLLM |
| Text to speech | Generic OpenAI-compatible, Speaches, Kokoro-FastAPI, vLLM-Omni, NeMo-Speech.cpp |
Choose the profile matching the server. Use Generic OpenAI-compatible for another compatible service, then check its requirements. Server version and model determine actual support; enabling a use does not verify it.
Decide where models run
Section titled “Decide where models run”| Location | Consider when… |
|---|---|
| This computer | You want inference to use this computer’s resources |
| Another computer or private server | You want to preserve local memory/GPU capacity or share a server |
| Hosted provider | You want a compatible service without managing its runtime |
A local GPU is optional. Network speed and server capacity affect response time. Audio or text goes to the endpoint you choose, whose data policy applies. Freehand is free forever; provider usage and server hosting costs are separate.
Choose a topology
Section titled “Choose a topology”Tasks can share a gateway or use different servers. Each task keeps its own model and options; tasks sharing a connection share its address and authentication.
Save one connection and enable the operations the gateway exposes. Select it independently for transcription, cleanup, or speech.
Transcription ─┐Cleanup ───────┼─ https://speech.example.com/v1Speech ────────┘A metadata check does not verify every enabled operation.
Save a connection for each server, then choose it in the relevant task.
Transcription ── Speech serviceCleanup ──────── Chat serviceSpeech ───────── Speech synthesis serviceA cleanup outage preserves the successful raw transcript for normal delivery or copying.
Local and LAN servers
Section titled “Local and LAN servers”For http://127.0.0.1:8000/v1 or a private address such as
http://192.168.1.50:8000/v1, explicitly enable Allow HTTP for this connection.
Freehand otherwise rejects plaintext requests.
Test without invoking a model
Section titled “Test without invoking a model”Choose Check connection using saved settings. The test reads a health route or model-list metadata; it never sends audio or a prompt or runs listed models.
| Result | Meaning |
|---|---|
| Model list received | Valid model inventory returned; review whether your selected ID is listed |
| Server reachable | The configured health endpoint responded successfully |
| Invalid response | The response was incompatible, including an HTML gateway page with status 200 |
| Selected model Not listed | Your ID was absent, including from an otherwise valid empty inventory |
Checks show HTTP status, latency, and a failure category. They cannot prove a model can transcribe or synthesize, or that an endpoint enforces the supplied key. See connection-check results.
Use a custom health path
The optional path is appended beneath the base URL path. A base URL of
https://speech.example.com/v1 plus /health checks
https://speech.example.com/v1/health, not the host-root route. Include the
leading slash. A failed custom check is reported without falling back to model
listing. Leave it blank for the backend default: /health for whisper.cpp or
model-list metadata for other backends.
Initial Voice setup requires Check connection before Finish setup. Afterward, the saved Voice connection is checked on launch and when relevant connection settings change. Correct a failed check and retry.
Choose a language
Section titled “Choose a language”Open the task’s Transcription → Spoken language and search by language or code. Server default, Automatic detection, named languages, and custom values depend on the selected model. This requests transcription in the source language, not translation.
Transcription controls
Section titled “Transcription controls”Save options before starting a new request. Voice shows context with model options and temperature under Request settings; Audio file groups them in Transcription controls. Only supported controls are offered.
| Control | Use |
|---|---|
| Transcription context | Short subject matter or expected wording; availability varies by backend, model, and mode |
| Vocabulary | Shared names and phrases, enabled separately for Voice and files; see Vocabulary |
| Override temperature | A value from 0 to 1 where supported; leave off for the server default |
Context and vocabulary guide recognition rather than editing its result. Context is limited to 8,192 UTF-8 bytes, including vocabulary when sent as context; your model may accept less. These hints are saved locally and sent to the selected server when used. Keep secrets out of them.
To compare an option, explicitly transcribe a short sample with your selected model, change one setting, and repeat. A rejected request is not retried with different options automatically. Use Cleanup to edit text.
Example: Speaches on your PC
Section titled “Example: Speaches on your PC”The Speaches launch guide covers GPU/CPU Docker setup, installing one chosen model, and connecting Freehand. For other deployments, follow the runtime’s own guide.