Skip to content

Nemotron 3.5 ASR streaming

The Nemotron 3.5 ASR streaming model profile supports NVIDIA’s Nemotron 3.5 ASR streaming 0.6B through NeMo-Speech.cpp. NeMo-Speech.cpp is the server; Nemotron is the model it runs.

For local use on Windows or macOS, follow managed NeMo setup. Nemotron is the recommended managed speech model. Freehand supplies its profile when you select the built-in Connection; you enable live mode in Voice.

For a manual Connection, choose Nemotron 3.5 ASR streaming beneath the model picker in Voice transcription → Transcription or Audio file → Transcription options.

Setting Completed recordings and audio files Realtime microphone
Spoken language Automatic or one of 32 base-model locales Same language choices
Shared vocabulary Recognition phrases with a strength setting Same vocabulary support
Realtime transcription Voice can switch to live mode Live text in the results pane
Live overlay captions Not used Optional single-row caption strip

Selecting this profile shows the supported language list and vocabulary controls. In Voice, it also reveals Realtime transcription inside the Transcription panel. Enabling it uses your selected connection and model; turning it off restores completed recording and checkpoint behavior. Audio-file settings remain independent.

For a manually configured NeMo service:

  1. Follow the NeMo-Speech.cpp setup guide to add the server connection.
  2. Open the workflow’s Settings cog and Transcription options, then select that connection and the loaded model ID.
  3. Select Nemotron 3.5 ASR streaming as the model profile, then choose automatic detection or your spoken language.
  4. For live dictation, enable Realtime transcription. Enable Live overlay captions for the caption strip; the main Overlay preference must also be on.
  5. Choose Save in the options sidebar.

Keep names and terminology in Settings → Vocabulary, then enable the list for Voice, audio files, or both. The recognition list has these limits:

Setting Limit
Phrases 32
Each phrase 128 UTF-8 bytes
Entire list 2,048 UTF-8 bytes
Vocabulary strength 0–5

Reuse the list with other supported models; see Vocabulary.

By default, Freehand requests verbatim output with the model’s native punctuation. NeMo transcription controls let you change punctuation and request server-side normalization or filtering. Vocabulary guides recognition; use the separate Cleanup stage for rewrite instructions.

How language evidence affects cleanup

Freehand removes the terminal language tag from displayed text and retains it as language evidence alongside structured metadata. Realtime finals retain languages across the entire recording: a later English turn does not erase earlier non-English evidence for cleanup.