Skip to content

Kokoro-FastAPI

Choose Kokoro-FastAPI for on-demand speech playback. The dedicated profile supports voice discovery, manual voice IDs, speed control, and buffered PCM16 WAV audio. It does not provide transcription or transcript cleanup.

  1. Open Connections, create a connection, and select Kokoro-FastAPI with Text to speech enabled.
  2. Enter the server’s API base URL, including /v1. For the local Docker example below, use http://127.0.0.1:8880/v1. For a remote server, use its HTTPS URL and the authentication supplied by its administrator.
  3. For the local HTTP example, enable Allow HTTP for this connection, then save. In Text to speech → Speech, select the connection and choose the model ID advertised by your server, normally kokoro.
  4. Click Refresh voices, then search or choose a voice such as af_heart. Manual IDs remain available, including administrator-provided aliases.
  5. Enable text to speech and preview the voice. Choose Save to apply your choices.

The voice is remembered per connection and model. Speaking speed stays the same when you switch either. Model and voice refreshes read metadata without synthesis.

These PowerShell examples require Docker Desktop using Linux containers. The CPU image is enough to get started and requires no dedicated GPU. Images include model assets; first use downloads the image. The upstream latest tag changes, so pin a release tag or image digest for a repeatable deployment.

Start Kokoro on CPU
docker run -d --name freehand-kokoro --restart unless-stopped `
-p 127.0.0.1:8880:8880 `
ghcr.io/remsky/kokoro-fastapi-cpu:latest

Choose one container recipe for this name and port. Inspect or control it with:

View output, stop, or restart
docker logs -f freehand-kokoro
docker stop freehand-kokoro
docker start freehand-kokoro

The bundled API documentation is at http://127.0.0.1:8880/docs. This recipe binds to loopback for a client on the same machine; see connection topologies for a separate server. Freehand does not start or manage the container.

Capability Request and response
Model discovery GET /v1/models; listed IDs may be compatibility aliases.
Voice discovery GET /v1/audio/voices; current ID/name objects and legacy strings.
Speech POST /v1/audio/speech with model, input, voice, speed, response_format: "wav", and stream: false.
Playback Fully buffered PCM16 WAV; speed requests from 0.25× through 4×.

Language overrides, server normalization, voice-blend creation, cloning inputs, captioned audio, and progressive playback are not available in this profile. Speech requests use your configured timeout and a 32 MiB response limit.