Skip to content

MagpieTTS Multilingual 357M

Freehand’s MagpieTTS Multilingual 357M profile supports the v2602 checkpoint through NeMo-Speech.cpp v0.1.0. It generates spoken audio from text with a choice of voices and languages.

Follow Local speech with MagpieTTS. NeMo can load Magpie alongside your transcription model. Both share the runtime’s Start and Stop controls.

Turn on Enable text to speech, choose a voice and language, and Save. Choose Preview to hear your current edits without saving them. Enter text in the composer and choose Speak to use the saved settings. These actions request audio; opening settings and refreshing metadata do not.

Choice Available values
Voices John, Sofia, Aria, Jason, Leo; each supports all checkpoint languages
Default voice default uses the server’s configured speaker
Languages English, Spanish, German, French, Italian, Vietnamese, Hindi, Mandarin Chinese, Japanese
Server default language Uses server configuration; does not request automatic detection

Refresh voices reads the loaded model’s metadata. After a successful refresh, Freehand limits language choices to those the server advertises. Japanese and Chinese require NeMo builds with their language frontends enabled.

Capability Availability
Voice and language Sent with the request
Audio Complete WAV output
Speed adjustment, style, cloning, streamed speech Unavailable
Text normalization Requires server support and grammar assets; managed setup does not install them

Your speech voice and language are saved with this model’s options. Generated audio stays in memory until cleared, replaced, recording starts, or Freehand quits. You can explicitly save a WAV file from the player.

Managed setup downloads the versioned Magpie model, NanoCodec decoder, and tokenizer assets qualified with NeMo v0.1.0.

See the v2602 model card, NeMo speech API, and language frontend requirements for the qualified server behavior.