Qwen3-ASR
Choose vLLM for the connection and Qwen3-ASR for the model profile. The model profile enables Qwen’s language and context controls. Select it yourself; Freehand does not infer it from a model name or download models.
Connect Freehand
Section titled “Connect Freehand”Have vLLM 0.28.0 running the Qwen3-ASR 1.7B checkpoint with the realtime architecture below. If you need to start it, follow Run the server first, then return here with its base URL and authentication details.
- Create a connection with backend vLLM, base URL such as
http://127.0.0.1:8089/v1, and Voice transcription or Audio-file transcription enabled under Used for. A local unauthenticated server uses None and requires Allow HTTP for this connection to be enabled. Remote deployments use their own address and authentication. - In Voice → Transcription, select that connection, check its model list, choose Qwen/Qwen3-ASR-1.7B, then choose the Qwen3-ASR model profile.
- Enable Realtime transcription inside this same panel for live results. Live overlay captions shows the latest text in the single-row overlay when the main Overlay preference is also on. Stop recording to finalize, optionally clean up, and safely insert.
- Turn realtime off for completed recordings and pause-aware checkpoints. Select the connection and model separately in Audio file to use files.
- Choose Save to apply the workflow settings.
Metadata checks only read the server’s health/model endpoints. These checks do not start transcription.
Supported controls
Section titled “Supported controls”| Setting | Completed recording / audio file | Realtime microphone |
|---|---|---|
| Model profile | Qwen3-ASR | Same profile and selected model |
| Language | Automatic or a supported language hint | Automatic only |
| Context hint | Recognition context, sent as prompt |
Unavailable |
| Shared Vocabulary | Appended to context when enabled | Unavailable; list and preference preserved |
| Temperature | Optional 0–1 request override | Unavailable |
| File response streaming | Supported | Separate WebSocket audio transport |
| Caption preview | After completed transcription | Provisional live text; never inserted directly |
Dialects, translation, forced alignment, timestamps, diarization, and model training have no controls in this profile.
Context and vocabulary are recognition hints, not a chat instruction or a guarantee of spelling. The shared library remains in Settings → Vocabulary. Completed settings stay saved when realtime is enabled and apply again when it is disabled. Realtime sends no language, prompt, vocabulary, or temperature.
Run the server
Section titled “Run the server”Use Linux vLLM 0.28.0 runtime with its audio dependencies. For Windows, Docker Desktop’s WSL2 GPU backend is one deployment option. The vLLM Docker setup documents the pinned image and required audio packages.
Serve the 1.7B checkpoint with the realtime architecture override. The same loaded model also serves completed audio; a second model is unnecessary. For example, in the Linux runtime:
vllm serve Qwen/Qwen3-ASR-1.7B \ --revision 7278e1e70fe206f11671096ffdd38061171dd6e5 \ --hf-overrides '{"architectures":["Qwen3ASRRealtimeGeneration"]}' \ --host 0.0.0.0 --port 8000 \ --gpu-memory-utilization 0.55 --max-model-len 4096 --max-num-seqs 1 \ --enforce-eager --no-enable-log-requests --disable-log-statsPublish the container port as 127.0.0.1:8089:8000 to connect from the same PC. A remotely accessible server needs the
deployment’s authentication and network protections. Keep request-content
logging disabled. No forced-aligner model is needed.
Check http://127.0.0.1:8089/health and
http://127.0.0.1:8089/v1/models after startup. Freehand derives
ws://127.0.0.1:8089/v1/realtime from the base URL; enter the HTTP base URL,
not the WebSocket URL, in Connections. An HTTPS base uses WSS automatically.
Stop a competing GPU inference model before starting this one if capacity is
limited. Model installation, startup, shutdown, and tuning remain server tasks.
Server version
Section titled “Server version”This profile uses vLLM 0.28.0 with the Qwen/Qwen3-ASR-1.7B checkpoint and
realtime architecture shown above. Use that architecture to enable realtime
alongside completed transcription.
Sources: Qwen model card, vLLM realtime protocol, Qwen realtime implementation, and completed transcription implementation.