Skip to content

NeMo-Speech.cpp

Freehand’s NeMo-Speech.cpp backend profile supports the v0.1.0 server API for microphone recordings, audio-file uploads, realtime microphone audio, and MagpieTTS speech generation.

Task Model profile
Completed or realtime transcription; language and vocabulary controls Nemotron 3.5 ASR streaming
Completed transcription with automatic language detection Parakeet TDT v3
Speech generation with voices and languages MagpieTTS

The server owns model loading, GPU selection, and chunk latency. It can run on the same computer or another machine reachable from Freehand.

NeMo can load one transcription engine and one speech-generation engine in the same process.

Shared runtime behavior What it means for you
Both enabled engines load before the listener is ready Allow memory and startup time for both models
One engine per capability The request’s model field does not switch among several transcription models
Shared process Starting, stopping, or restarting NeMo affects transcription and speech together

Managed setup lets you select Nemotron or Parakeet plus optional MagpieTTS. For a manual server, set the model paths in YAML configuration and follow engine and listener configuration. Freehand’s Connections select how tasks use that server; they do not load extra models on a remote server.

  1. Open Connections and choose Add connection.
  2. Choose NeMo-Speech.cpp, name the connection, and enable the uses you need: Voice transcription, Audio-file transcription, and Text to speech.
  3. Enter the HTTP API root, such as http://127.0.0.1:8088/v1 for a local server on port 8088. Enable Allow HTTP for this connection when using HTTP on a trusted network; add authentication if required by the server.
  4. Save the connection, then open the Settings cog in a workflow and select it in Transcription or Speech options.
  5. Check the connection and select the model loaded on the server. Choose Nemotron 3.5 ASR streaming or Parakeet TDT v3 for transcription, and MagpieTTS Multilingual 357M for speech generation. Save the workflow settings.

Connection checks read metadata only. They show advertised capabilities and device, plus the version when /health is available. The qualified version is 0.1.0; metadata does not prove inference support or choose a model profile.

Which models appear in each workflow?

Each workflow filters models by advertised capability. Transcription also keeps entries whose capability is unknown. Connection diagnostics can show both transcription and speech models from a combined server. Generic model behavior offers completed transcription; choose Nemotron explicitly for Voice’s Realtime transcription switch.

Operation API and result
Completed recordings and files /v1/audio/transcriptions with verbose_json for text, language, and duration; no partial file-response streaming
Realtime microphone /v1/audio/transcriptions/realtime; the NeMo transcription WebSocket, without fallback to a general realtime route
Speech /v1/audio/speech, buffered WAV, with MagpieTTS voice and language controls
Refresh voices Reads the speech model’s voices and languages in /v1/models; does not synthesize audio

Enter the HTTP base URL. Freehand derives the WebSocket address; HTTPS uses WSS, and a configured API key goes in the upgrade header. Turning realtime off keeps the same Voice connection and model for completed recordings.

Choose the qualified Nemotron 3.5 ASR streaming or Parakeet TDT v3 model profile, then open NeMo transcription controls in the workflow’s Transcription options. Voice and audio files remember their controls separately for each connection and model. Save applies changes to the next request.

Control Default Effect and prerequisites
Automatic punctuation On Keeps model punctuation and server formatting. Off asks NeMo to strip formatting.
Normalize numbers and dates Off Requests inverse text normalization (ITN), such as “twenty one” → “21”. Requires a server build with ITN support and configured language grammars.
Profanity filter Off Requests masking using a word list configured on the server.
Server endpointing delay (ms) 0: server default Realtime only. A value from 100 to 10,000 adjusts the pause threshold when server endpointing is already enabled. It does not enable endpointing or change Freehand’s microphone stop controls.

These controls follow the pinned NeMo server API and ASR configuration.