Skip to content

Voxtral Mini Realtime

Use mistralai/Voxtral-Mini-4B-Realtime-2602 with vLLM v0.28.0 and Freehand’s Voxtral Mini Realtime model profile.

Set up vLLM on your inference machine with its audio dependencies, then serve the checkpoint:

Serve Voxtral Mini Realtime
vllm serve mistralai/Voxtral-Mini-4B-Realtime-2602
  1. Add a vLLM connection using the server’s HTTP API root, ending in /v1.
  2. In Voice transcription, choose the connection and the served model ID.
  3. Choose Voxtral Mini Realtime as the model profile.
  4. Enable Realtime transcription and optionally Live overlay captions. Captions also require the main Overlay preference to be on. Save your changes.

Freehand derives /v1/realtime from the same connection. The recording shortcut starts microphone streaming; provisional words appear in Current result and, when enabled, the single-row overlay caption. Stop recording to receive the final text and apply your usual cleanup and focus-safe insertion.

Mode or setting Behavior
Realtime on Live microphone previews, followed by final text when recording stops
Realtime off Same model for completed recordings and checkpoints
Audio-file transcription Independent selection; uploads return completed text
Language Automatic detection
Context, vocabulary, temperature No controls in this profile

See vLLM’s realtime example for the server’s microphone protocol.