NeMo-Speech.cpp
Freehand’s NeMo-Speech.cpp backend profile supports the v0.1.0 server API for microphone recordings, audio-file uploads, realtime microphone audio, and MagpieTTS speech generation.
| Task | Model profile |
|---|---|
| Completed or realtime transcription; language and vocabulary controls | Nemotron 3.5 ASR streaming |
| Completed transcription with automatic language detection | Parakeet TDT v3 |
| Speech generation with voices and languages | MagpieTTS |
Run the server
Section titled “Run the server”The server owns model loading, GPU selection, and chunk latency. It can run on the same computer or another machine reachable from Freehand.
Combined transcription and speech
Section titled “Combined transcription and speech”NeMo can load one transcription engine and one speech-generation engine in the same process.
| Shared runtime behavior | What it means for you |
|---|---|
| Both enabled engines load before the listener is ready | Allow memory and startup time for both models |
| One engine per capability | The request’s model field does not switch among several transcription models |
| Shared process | Starting, stopping, or restarting NeMo affects transcription and speech together |
Managed setup lets you select Nemotron or Parakeet plus optional MagpieTTS. For a manual server, set the model paths in YAML configuration and follow engine and listener configuration. Freehand’s Connections select how tasks use that server; they do not load extra models on a remote server.
Connect Freehand
Section titled “Connect Freehand”- Open Connections and choose Add connection.
- Choose NeMo-Speech.cpp, name the connection, and enable the uses you need: Voice transcription, Audio-file transcription, and Text to speech.
- Enter the HTTP API root, such as
http://127.0.0.1:8088/v1for a local server on port 8088. Enable Allow HTTP for this connection when using HTTP on a trusted network; add authentication if required by the server. - Save the connection, then open the Settings cog in a workflow and select it in Transcription or Speech options.
- Check the connection and select the model loaded on the server. Choose Nemotron 3.5 ASR streaming or Parakeet TDT v3 for transcription, and MagpieTTS Multilingual 357M for speech generation. Save the workflow settings.
Connection checks read metadata only. They show advertised capabilities and
device, plus the version when /health is available. The qualified version is
0.1.0; metadata does not prove inference support or choose a model profile.
Which models appear in each workflow?
Each workflow filters models by advertised capability. Transcription also keeps entries whose capability is unknown. Connection diagnostics can show both transcription and speech models from a combined server. Generic model behavior offers completed transcription; choose Nemotron explicitly for Voice’s Realtime transcription switch.
Supported API
Section titled “Supported API”| Operation | API and result |
|---|---|
| Completed recordings and files | /v1/audio/transcriptions with verbose_json for text, language, and duration; no partial file-response streaming |
| Realtime microphone | /v1/audio/transcriptions/realtime; the NeMo transcription WebSocket, without fallback to a general realtime route |
| Speech | /v1/audio/speech, buffered WAV, with MagpieTTS voice and language controls |
| Refresh voices | Reads the speech model’s voices and languages in /v1/models; does not synthesize audio |
Enter the HTTP base URL. Freehand derives the WebSocket address; HTTPS uses WSS, and a configured API key goes in the upgrade header. Turning realtime off keeps the same Voice connection and model for completed recordings.
Transcription controls
Section titled “Transcription controls”Choose the qualified Nemotron 3.5 ASR streaming or Parakeet TDT v3 model profile, then open NeMo transcription controls in the workflow’s Transcription options. Voice and audio files remember their controls separately for each connection and model. Save applies changes to the next request.
| Control | Default | Effect and prerequisites |
|---|---|---|
| Automatic punctuation | On | Keeps model punctuation and server formatting. Off asks NeMo to strip formatting. |
| Normalize numbers and dates | Off | Requests inverse text normalization (ITN), such as “twenty one” → “21”. Requires a server build with ITN support and configured language grammars. |
| Profanity filter | Off | Requests masking using a word list configured on the server. |
| Server endpointing delay (ms) | 0: server default | Realtime only. A value from 100 to 10,000 adjusts the pause threshold when server endpointing is already enabled. It does not enable endpointing or change Freehand’s microphone stop controls. |
These controls follow the pinned NeMo server API and ASR configuration.