vLLM-Omni
Freehand supports vLLM-Omni v0.18.0 for buffered WAV speech generation. Choose this backend separately from the vLLM transcription or cleanup backend. For text to speech, select vLLM-Omni in Connections.
Run the server
Section titled “Run the server”Follow the runtime’s installation and Qwen3-TTS serving guide on your inference machine. Load Qwen/Qwen3-TTS-12Hz-1.7B-CustomVoice and expose its HTTP API. Configure model loading and the GPU on that server.
Connect Freehand
Section titled “Connect Freehand”- Open Connections → Add connection.
- Choose vLLM-Omni, select Text to speech under Used for, and enter the HTTP API root,
such as
http://127.0.0.1:8091/v1for a local service on port 8091. - Configure authentication. If using HTTP on a trusted network, enable Allow HTTP for this connection, then save.
- Open Text to speech and its Settings cog, then select the connection and model in Speech options.
- Choose the Qwen3-TTS model profile to use its preset speakers, ten languages, and voice-style instructions.
- Turn on Enable text to speech, choose a preset voice, and use Preview to hear it. Choose Save.
Supported API
Section titled “Supported API”| Operation | Request or result |
|---|---|
| Connection check | Model metadata only |
| Refresh voices | GET /v1/audio/voices; no synthesis |
| Speech generation | POST /v1/audio/speech, response_format: "wav", stream: false, then native playback |
| Qwen CustomVoice controls | Adds task_type: "CustomVoice", language name, and optional instructions |
Generic model behavior offers standard voice-ID and speed fields when the loaded model accepts them. Choose the Qwen profile only for its matching checkpoint. Reference-voice uploads and voice-cloning tasks are unavailable.
Choose Qwen3-TTS controlsPreset speakers, supported languages, and voice-style instructions.