Full-duplex voice mode for the DeepSeek Harness Web UI: toggle (2s-pause auto-send) or hold-to-talk dictation into an editable draft with zipformer2 streaming ASR, optional wake word; sentence-by-sentence Edge TTS read-aloud with live captions, and speaking interrupts playback and the running turn (true barge-in). On-device ASR, no API key.
Install
# from npm (prebuilt)
dsh plugin --profile web add dsh-voice-mode
# from GitHub (first run asks for allowBuilds approval — follow the hint, retry)
dsh plugin --profile web add github:qishuilalala/dsh-voice-mode#path:/plugin/dsh-voice-mode
Any plugin you install runs third-party code with your own permissions — it can read your files, use your credentials, and reach the network, and tool approvals don’t sandbox it. GitHub-sourced plugins also run build scripts at install time — pnpm blocks those until you allow them, so an install can stop with ERR_PNPM_GIT_DEP_PREPARE_NOT_ALLOWED or ERR_PNPM_IGNORED_BUILDS; dsh prints the exact key to add under allowBuilds in your profile’s pnpm-workspace.yaml, and the install works on the next run. Allowing a build is a trust decision: only install sources you trust, and pin a commit (github:owner/repo#sha).
README
Full-duplex voice conversation mode for DeepSeek Harness (dsh): speak, get a spoken answer. Streamed zipformer2 ASR → editable draft → auto send → the final reply is read out sentence-by-sentence via Edge TTS, and your voice interrupts playback and the running turn. No API key.
English · 中文


Runtime changes (most recent batch: v0.7.17 – v0.7.19, 2026-10-02): (1) Settings work on every dsh version — fixed the settings page and persistence on dsh 0.1.7+ (dedicated Settings → Voice Mode page, writes survive restarts); (2) i18n reworked — the UI follows dsh's official language setting, Chinese/English are fully consistent and switch without a reload; (3) fixed fresh-install failures with npm/pnpm and
failed to importon dsh ≥0.1.7 (affects 0.7.17 and earlier — please upgrade to ≥0.7.18); (4) local Kokoro fixed on the official Electron desktop host — needs Node.js ≥18 on PATH (see the desktop note below). Earlier batches (v0.7.10: silent-audio fixes, etc.) and the full history are inCHANGELOG.md.
🤝 Fork enhancements (this repo)
- Edge cloud TTS by default; local TTS optional (privacy-first): local VITS (Chinese) and local Kokoro (zh+en, 103 voices;
int8default ~109 MB /fp32~311 MB for better quality; nativesherpa-onnx-nodeaddon, no WASM memory limits) run in an isolated child process; - 103 Kokoro voices (F0-measured gender labels, 4 favourite male voices pinned), browsed with a ◀▶ stepper;
- Delta transport: partials upload only the new 0.9 s — long push-to-talk segments finalize in seconds;
- Interaction: a mode-switch button next to the mic (continuous ⇄ hold, persisted); hold mode records only while held;
- Long segments: continuous listening stitches consecutive segments into one message (internally chunked at 30 s and concatenated across chunks); 1500 ms silence split by default; hold keeps pauses from splitting;
- Hardening: session-existence check, loopback + Origin guards, per-endpoint rate limits, model SHA256 pinning, download-host allowlist.
ℹ️ The
demo-voice-flow.gifat the top is a real recording of the current UI (voice mode → live transcription → pause auto-send → sentence-by-sentence read-aloud with live captions), produced by driving the real pipeline viascreenshots/scripts/capture-demo.mjs.
✨ Features
- Voice mode: toggle with the microphone button in the input toolbar or the global shortcut
Ctrl+Shift+V; globally single-active (only one session is in voice mode at a time; switching sessions yields automatically) - Two interaction modes (switchable in settings, plus a mode-switch button beside the mic):
toggle(default) continuous listening: RMS VAD segmentation → streaming zipformer2 ASR (words appear as you speak, live caption preview) → automatic sentence split after 1500 ms of silence into the draft, consecutive segments joined into one message, then auto-sent after ~1500 ms more of silence (≈3 s total); holdCtrlto force an immediate sendholdpush-to-talk: short tap to enter/exit, hold the mic button to talk, release to send (swipe up to cancel,Esc/blur abandons the segment; pauses do not split while held, up to 10 min); holdCtrlto record-by-keyboard, release to send
- Wake word (optional, off by default): after setting
wakeWord, entering voice mode starts in standby, and recognition only begins once the wake word is spoken (e.g.你好小D). You may say it together with your command ("wake word, check the weather") — the wake word is stripped from the transcript/caption and never sent; matching is fault-tolerant (homophone substitutions, a wrong first char, leading fillers); 3-4 characters recommended; in standby the status bar live-shows "say '' · what it heard". Boundaries (important): toggle mode only (inactive withhold/ manual barge-in); it must be repeated after each completed utterance or barge-in (if you only said the wake word, the engine keeps listening so you may pause before the command); saying it while the agent is reading does not trigger (barge-in is VAD-based and unrelated to the wake word) - Caption tiers (batch 3):
captionFontSize4 levels (0=12px / 1=14px / 2=18px / 3=24px) +captionMaxWidth3 levels (0=50vw / 1=70vw / 2=90vw); adapts to narrow viewports - Yield semantics (batch 5 / ADR-0008):
backchannelYieldon by default (I10-exempt) — saying嗯 / 对 / rightwhile the AI is reading auto-pauses for 1.5 s; genuine speech still triggers hard barge-in; lets the LLM yield the turn - Output pipeline: only the final answer's
text-deltais read (reasoning/tool calls are skipped), streamed sentence-by-sentence (Edge cloud by default; local VITS / Kokoro, int8/fp32, optional) with a live caption overlay at the bottom-right; tool calls trigger a beep; the full text is still written to the chat; in voice mode a spoken-format system prompt is injected (short natural sentences, no Markdown decoration), and the reader side strips markers as well for a smoother listening experience - Barge-in: three sensitivity levels of voice-onset detection → local mute + host synth queue invalidation (epoch) + running turn cancellation (the half-finished part is kept and naturally flows into your new message). With a wake word set, barge-in remains speak-to-interrupt (the gate is the VAD, not the wake word; the engine returns to standby afterwards)
- Lazy model download with progress: the zipformer2 Chinese streaming model (~160 MB,
.partresumable) is downloaded on first use with live progress in the status bar;npm run prefetchcan pre-download it - Resilience: mic-denied red hint, visible model-download failure, TTS unreachable status hint (auto retry), failed submit keeps the text in the draft, SSE auto-reconnect
- Settings: location depends on your dsh version (see Settings below), with voice / rate / interrupt sensitivity / silence pause / idle timeout / model mirror / auto send / interaction mode / wake word / ITN / caption tiers / yield semantics; voices are previewable (the "试听/Preview" button synthesizes and plays the current voice at the current rate instantly, no need to enter voice mode; custom ShortNames are previewable too)
- Interface language: follows dsh's language setting (Settings → General → Language) and switches instantly, no reload — the settings page, voice status bar, captions, voice names/gender/accent and error messages all follow it, and the LLM spoken-style prompt is chosen in English or Chinese accordingly. Dictionaries are registered in dsh's official locale service under the
voice-modenamespace, so community language packs can add more languages (missing ones fall back to English). The plugin's name and description in the Plugins page are localized too. - Idle exit: auto-exit and mic release after 5 minutes of inactivity (reading counts as activity; batch J 10→5)
⌨️ Interaction gestures
| Gesture | Behaviour |
|---|---|
Click the mic button / Ctrl+Shift+V |
Enter / exit voice mode |
| Just speak, pause ~1500 ms (toggle) | Accumulate into draft, auto-send after ~3 s of quiet |
Hold Ctrl (toggle, ≥250 ms speech) |
Force-send the current segment immediately |
| Hold the mic button (hold) | Hold to talk, release to send; swipe up / Esc / blur abandons the segment; <250 ms tap exits the mode |
Hold Ctrl (hold, ≥600 ms) |
Keyboard push-to-talk, release to send |
| Speak the wake word first (if configured) | Activate from standby into listening (then recognition and sending begin) |
| Speak while AI is reading | Interrupt playback and cancel the running turn |
| Type in the input box | Auto-exit voice mode (draft is kept) |
🚀 Quick Start (5 minutes)
Requirements: dsh web (Node ≥ 18), a modern browser (Chrome / Edge / Firefox, supporting getUserMedia and Web Audio).
# Option 1: from npm (recommended)
dsh plugin --profile web add dsh-voice-mode
# Equivalent via npx if the dsh CLI is not installed locally:
npx -y @deepseek-ai/dsh plugin --profile web add dsh-voice-mode
# Option 2: local tarball
dsh plugin --profile web add ./dsh-voice-mode-0.1.0.tgz
# Option 3: from source
git clone https://github.com/qishuilalala/dsh-voice-mode.git
cd dsh-voice-mode/plugin/dsh-voice-mode && pnpm install && pnpm build
dsh plugin --profile web add .
Bundle plugins require a dsh restart to take effect (restart methods by platform):
- Linux (systemd):
systemctl restart dsh - Windows / macOS / manual hosting: restart your dsh process (kill and run
dsh webagain, or restart it in your service manager)
Optional: pre-download the ASR model to reduce the download wait on the first voice-mode entry:
npm run prefetch # run inside the plugin dir; writes to the platform cache dir
# or specify the cache location: node scripts/prefetch.mjs --cache-dir /where/ever/models
First run:
- Click the mic button in the input toolbar (or press
Ctrl+Shift+V) to enter voice mode; a status bar appears above the input box - Choose how to speak: just talk and let the ~1500 ms pause split and ~3 s of quiet auto-send (toggle); or hold the mic button and release to send (hold)
- The AI answer is read sentence-by-sentence with a caption overlay at the bottom-right; click "Skip" or just start speaking to interrupt
- Click "Exit" in the status bar (or press
Ctrl+Shift+Vagain) to leave voice mode
On first entry the recognition model is downloaded; the status bar shows Loading model… <file> <percent>%.
If a wake word is configured, you land in standby first (the status bar prompts Say "<wake word>" to start), and recognizing starts after you speak the wake word.
⚙️ Settings
Works on every dsh version: Settings → Voice Mode (the plugin registers its own page in the Settings dialog). Secondary entry points (same card) vary by version:
- dsh ≤ 0.1.5: Settings → Plugins → plugin config → voice mode
- dsh 0.1.6-alpha and later (incl. 0.1.7 / 0.2.0): left sidebar Plugins → the installed
dsh-voice-mode→ the "Voice Mode" card on its detail page - On dsh 0.1.7+ the values are stored in the plugin's own file
~/.dsh/voice-mode.settings.json(under$DSH_HOME) and take precedence over same-named keys in the profile config. The first run migrates the oldvoice-mode:block fromsettings.yaml(.imported)once. To reset, set the file's content to{}(do not delete it).
4 new settings (11 batches of comprehensive fixes)
| Key | Default | Description |
|---|---|---|
senseITN |
true |
Batch 2 P0: SenseVoice inverse text normalization (number / date / currency; on by default) |
captionFontSize |
0 |
Batch 3 P0: caption font tier 0=12px / 1=14px / 2=18px / 3=24px (default 0 is byte-equivalent to legacy) |
captionMaxWidth |
1 |
Batch 3 P0: caption width tier 0=50vw / 1=70vw / 2=90vw |
backchannelYield |
true |
Batch 5 P1: yield semantics (ADR-0008); saying 嗯 / 对 while reading auto-pauses for 1.5 s; genuine speech still triggers hard barge-in. I10-exempt (default-on is a product decision); off = behavior identical to pre-change |
5 defaults micro-adjusted (batch J)
| Key | Old | New | Why |
|---|---|---|---|
rate |
1.0 | 1.1 | Edge TTS defaults slightly slow; +10% improves perceived quality |
idleTimeoutMinutes |
10 | 5 | More responsive idle exit (reading still counts as activity) |
interruptLevel description |
old wording | new wording | Make "3/2/1 frame confirmation" explicit |
Field names unchanged → 100% backward compatible with existing
~/.dsh/settings.yaml.
Full 19-key settings table
| Key | Default | Description |
|---|---|---|
ttsEngine |
edge |
Read-aloud engine: edge Microsoft cloud (default, fast) / vits local Chinese / kokoro local zh+en; applies live |
kokoroModel |
int8 |
Kokoro precision: int8 (default, 109 MB, CPU/low-bandwidth) / fp32 (311 MB, better quality, GPU/large memory); same 103 voices; applies live |
voice |
per engine | Voice: 5 VITS speakers; 103 Kokoro voices (◀▶ stepper; 62/68/75/76 favourite males pinned); Edge ShortNames below. The inline "试听" button previews it at the current rate |
rate |
1.1 |
Reading speed multiplier (0.5 slow ~ 2.0 fast), applies live (batch J 1.0→1.1) |
interruptLevel |
0 |
Barge-in sensitivity (host-side VAD frame detection + echo gate): 0 high threshold (3 frames) / 1 medium (2 frames) / 2 low (1 frame) |
bargeInMode |
detect |
Barge-in mode: detect auto-probes native echo cancellation (default; falls back to hold-to-talk when it is not in effect) / auto force interrupt as soon as you speak (headphones / quiet rooms) / manual hold-to-talk (recommended on loudspeakers). Batch 7O default, I10-exempt (ADR-0006) |
echoGateDb |
6 |
Echo-gate threshold (dB): auto barge-in requires the residual to exceed the echo floor by this much. Idle while native AEC is in effect (it only acts as a fallback on Safari / headsets without native AEC); lower by 3-4 if you cannot interrupt, raise by 8-10 if noise interrupts. Do not tune it for "cannot interrupt" (see ADR-0006) |
silenceMs |
1500 |
Silence pause in ms that marks the end of a complete sentence |
idleTimeoutMinutes |
5 |
Minutes of inactivity before auto-exiting voice mode (reading counts as activity; batch J 10→5) |
modelHost |
default | Model download host (use https://hf-mirror.com on mainland networks) |
autoSend |
true |
Auto-send once quiet (consecutive segments join into one message); when off, text only goes to the draft (hold Ctrl / release in hold mode still sends) |
mode |
toggle |
Interaction mode: toggle continuous listening + 1500 ms silence split; hold push-to-talk, release to send (short tap exits) |
autoResume |
false |
Auto-restore voice mode when switching back to your last voice session (off by default). When on: the next time you enter that session it auto-enters voice mode and restores it; when off, press Ctrl+Shift+V manually |
shortcut |
Ctrl+Shift+V |
Shortcut to enter / exit voice mode (modifiers Ctrl/Shift/Alt/Meta + one letter key); empty = disabled, use the mic button instead |
wakeWord |
empty (off) | Wake word (e.g. 你好小D): speak it after entering to activate; empty = off. May be said together with your command — the wake word is stripped and never sent; fault-tolerant matching (edit distance ≤1 + a 3-char leading window absorbs homophones/fillers); 3-4 characters recommended (single-char words get no tolerance; a 2-char word also absorbs any same-first-char 2-char word); must be repeated after each utterance split or barge-in; toggle mode only; not triggered while the agent is reading |
spokenFormat |
true |
Spoken-format system prompt: inject "short natural sentences, no Markdown decoration" into voice-mode replies only, applies live |
senseVoice |
true |
Re-transcribe the finalized utterance with SenseVoice (punctuation + number normalization, more accurate; on by default). Turning it off saves ~228 MB of models and uses streaming recognition only (faster, less accurate) |
toolBeep |
false |
Tool-call beep (off by default): when on, a short beep plays each time the agent calls a tool; off = silent |
senseITN |
true |
Batch 2 P0: SenseVoice inverse text normalization (number / date / currency; on by default) |
captionFontSize |
0 |
Batch 3 P0: caption font tier 0=12px / 1=14px / 2=18px / 3=24px (default 0 is byte-equivalent to legacy) |
captionMaxWidth |
1 |
Batch 3 P0: caption width tier 0=50vw / 1=70vw / 2=90vw |
backchannelYield |
true |
Batch 5 P1: yield semantics (ADR-0008); saying 嗯 / 对 while reading auto-pauses for 1.5 s; genuine speech still triggers hard barge-in. I10-exempt (default-on is a product decision); off = behavior identical to pre-change |
yieldMs |
1500 |
Yield-window length (ms, 500-3000): how long TTS frames are dropped after backchannelYield triggers; if you really want to speak in the window the normal hardBreak takes over, and playback resumes when the window ends |
Effect timing: voice/rate/ttsEngine/kokoroModel/spokenFormat take effect immediately (TTS hot-swap); the rest apply on the next voice-mode entry. Defaults come from the plugin config (base layer) — they follow the config unless explicitly changed.
Common voices (full list: node scripts/list-voices.mjs)
Desktop (Electron) note: the official desktop host runs inside Electron, whose V8 forbids the external buffers returned by Kokoro's native add-on. The plugin automatically runs Kokoro in a real Node.js instead — it needs Node.js >= 18 on PATH (or set
DSHVM_NODEto its path). Without it Kokoro reports an error; use the default Edge engine or local VITS (both work on desktop).
| ShortName | Description |
|---|---|
zh-CN-XiaoxiaoNeural |
Xiaoxiao · Female (default) |
zh-CN-XiaoyiNeural |
Xiaoyi · Female |
zh-CN-YunxiNeural |
Yunxi · Male |
zh-CN-YunjianNeural |
Yunjian · Male |
zh-CN-YunyangNeural |
Yunyang · Male |
zh-CN-YunxiaNeural |
Yunxia · Male |
zh-CN-liaoning-XiaobeiNeural |
Xiaobei · Northeastern Mandarin · Female |
zh-CN-shaanxi-XiaoniNeural |
Xiaoni · Shaanxi Mandarin · Female |
zh-HK-HiuMaanNeural |
HiuMaan · Cantonese · Female |
zh-HK-WanLungNeural |
WanLung · Cantonese · Male |
zh-TW-HsiaoYuNeural |
HsiaoYu · Taiwanese Mandarin · Female |
zh-TW-YunJheNeural |
YunJhe · Taiwanese Mandarin · Male |
en-US-AriaNeural |
Aria · English · Female |
en-US-GuyNeural |
Guy · English · Male |
🔧 Configuration (bundle patch / settings.yaml)
You can also edit the voice-mode: section of ~/.dsh/settings.yaml directly (the GUI card and RPC write to the same document layer):
- id: voice-mode
name: dsh-voice-mode
config:
enabled: true # false = disables voice mode entirely (toggle rejects)
cacheDir: ~/.cache/dsh-voice-mode/models # overridable; platform default otherwise
# Defaults seeded for the settings (the settings panel overrides; the panel is authoritative):
voice: zh-CN-XiaoxiaoNeural
rate: 1.1 # batch J 1.0→1.1
interruptLevel: 0
silenceMs: 1500
idleTimeoutMinutes: 5 # batch J 10→5
modelHost: https://huggingface.co
Note: the effective values of
voice/rate/interruptLevel/silenceMs/idleTimeoutMinutes/modelHost/autoSendcome from the settings panel; the bundle config only seeds defaults for those keys (enabled/cacheDirremain bundle-config-only). The plugin HTTP namespace is fixed to/voice-mode(matching the client bundle contract; not configurable).
🌐 API
| Route | Description |
|---|---|
GET /voice-mode/stream |
SSE: event: audio ({sessionId, seq, text, audio(base64 MP3)}), event: mode (global single-active ownership), event: tool (beep), event: asr-progress / asr-ready / asr-error / tts-error |
POST /voice-mode/toggle |
{sessionId, on} enter/exit voice mode (globally single-active) |
POST /voice-mode/asr |
Raw f32 LE 16k PCM payload → {text} (streaming zipformer2); returns 202 {loading} until the model is ready; ?reset=1 discards the in-flight segment (used on wake-word hit) |
POST /voice-mode/cancel |
{sessionId} invalidates the TTS queue and drops the in-flight ASR segment |
POST /voice-mode/preview |
{voice, rate?} one-shot synthesis preview → audio/mpeg (400 missing voice / voice too long; 502 synthesis failure, e.g. invalid ShortName; 403 when the plugin's enabled=false). Does not require voice mode to be active; uses an isolated synthesis connection and does not affect the reading queue |
GET /voice-mode/config |
Client bootstrap parameters (silence threshold / sensitivity / voice and rate, etc.) — includes the ASR-side fields senseITN / senseVoice / captionFontSize / captionMaxWidth / backchannelYield |
GET /voice-mode |
Health check {ok, name, enabled, active} |
💾 Model & cache
- Recognition model:
csukuangfj/sherpa-onnx-streaming-zipformer-zh-int8-2025-06-30(encoder ≈154 MB / decoder / joiner / tokens, ~160 MB total), running host-side via sherpa-onnx (Node WASM, Apache-2.0, natively cross-platform) - Cache directory defaults by platform:
- Windows:
%LOCALAPPDATA%\dsh-voice-mode\models - macOS / Linux:
~/.cache/dsh-voice-mode/models - both overridable via
cacheDir
- Windows:
- Downloads use
.partresume;huggingface.cofalls back tohf-mirror.comon failure (configurable viamodelHost)
🏛️ How it works
flowchart LR
subgraph Client[Browser Client]
Mic[Microphone 16kHz<br/>AudioWorklet<br/>echoCancellation:true] --> VAD[Client-side VAD<br/>RMS segmentation]
VAD -->|partial 0.9s| UI[Statusbar + caption overlay]
end
subgraph Host[dsh.host]
ASR[zipformer2 streaming ASR<br/>host-side WASM]
SV[SenseVoice finalization<br/>+ ITN + punctuation]
Tap[llm/stream tap<br/>observe-only]
Seg[sentence segmenter]
Q[TtsQueue<br/>epoch barge-in]
TTS{Engine}
Edge[Edge cloud]
Vits[Local VITS WASM]
Kokoro[Local Kokoro<br/>native addon]
end
UI -->|audio f32 PCM<br/>POST /voice-mode/asr| ASR
ASR --> SV
SV --> Draft[composer draft<br/>autoSend]
Draft --> Tap
Tap --> Seg
Seg --> Q
Q --> TTS
TTS -->|edge| Edge
TTS -->|vits| Vits
TTS -->|kokoro| Kokoro
Edge -.->|SSE audio frame| UI
Vits -.->|SSE audio frame| UI
Kokoro -.->|SSE audio frame| UI
VAD -.->|wake-word / barge-in| Q
- Speech and reading only happen for the session pointed to by the global single-active pointer
activeVoiceSession; other sessions pass throughllm/streamwith zero overhead (mode isolation) - The
llm/streamtap is lossless: every chunk passes through unchanged; segmentation/synthesis only observe and never block the model stream - ASR runs host-side (sherpa-onnx WASM, zipformer2 Chinese streaming + SenseVoice finalization which adds punctuation); the browser only captures audio (
getUserMedia16k mono) and does endpoint detection - Local TTS (VITS / native Kokoro) runs in an isolated child process (fork, auto-restart); barge-in kills the in-flight synthesis instantly to free CPU
- The TTS queue is per-session with an epoch version: old frames are all invalidated after a barge-in, so it is truly silent
🔍 Comparison with dsh built-in voice mode
| Dimension | dsh built-in | dsh-voice-mode (this plugin) |
|---|---|---|
| Recognition model | Cloud API (needs key) | Local zipformer2 + SenseVoice (zero key) |
| Multi-language | English-first | SenseVoice auto-recognition + ITN |
| TTS engine | Cloud TTS | Edge cloud + local VITS/Kokoro (three-way switch) |
| Barge-in detection | Basic VAD | 3 sensitivity levels + echo gate + yield semantics |
| Hotword biasing | None | None (removed in v0.7.7; see version note) |
| Caption a11y | None | 4 font tiers + 3 width tiers + theme-following |
| Wake word | None | Lightweight streaming match + prefix filler whitelist |
| dsh compatibility | — | full range since 0.1.1-rc.2 (incl. 0.1.7-rc.1) |
🔐 Permissions and data flow
An honest disclosure of what the plugin does (mapped to the awesome-dsh-plugin capability scan: network / fs-read / fs-write / shell / env):
| Category | What it does | Where data goes |
|---|---|---|
| network | (1) The default read-aloud engine Edge sends the reply text to be spoken over WebSocket to Microsoft speech.platform.bing.com (cloud synthesis — switch to local VITS / Kokoro to stay fully offline); (2) downloads local models on first use from huggingface.co (or your configured mirror hf-mirror.com) with a host allowlist and pinned SHA256; (3) registers loopback HTTP routes /voice-mode/* on the dsh host (loopback only by default; LAN needs allowLan) |
Microphone audio never leaves your machine (recognition runs locally via sherpa-onnx); only reply text goes to Microsoft, and only with the Edge engine |
| fs-read / fs-write | Model cache (Linux/macOS ~/.cache/dsh-voice-mode/models/, Windows %LOCALAPPDATA%\dsh-voice-mode\models\); the dsh 0.1.7+ settings overlay $DSH_HOME/voice-mode.settings.json; never touches your workspace files |
All local |
| shell (child processes) | Local TTS synthesis runs in a separate child process (child_process.fork → lib/tts-vits-worker.cjs); when the host is Electron (the official desktop app), it additionally probes and launches a real Node.js for Kokoro (node -p … to check the version). It never executes user input and never builds shell command strings |
Local |
| env | Reads only DSH_HOME, LOCALAPPDATA (for directories) and the optional DSHVM_NODE; no API key is read or needed |
— |
Also: the on-device recording fixture used for debugging is off by default; the plugin sends no telemetry. See SECURITY.md for the threat model and how to report vulnerabilities.
Third-party models and services
- Local models are not shipped in the package: the ASR (streaming zipformer, SenseVoice re-transcription), VAD and local TTS (VITS / Kokoro) models are downloaded on first use from Hugging Face (or your configured mirror) into a local cache; they come from export repositories in the sherpa-onnx ecosystem (
csukuangfj/*). Each model is licensed by its original authors — this plugin's MIT license does not cover them; verify the upstream licenses before commercial use (the model list is under "Model & cache"; the constants are insrc/asr-host.ts/src/tts-local.ts). - The Edge cloud read-aloud is a third-party online service: the default engine uses the Microsoft Edge browser "Read Aloud" endpoint via
msedge-tts. It is not an officially supported Microsoft API — availability, rate limits and terms are Microsoft's and may change at any time. For stability/compliance-sensitive use, switch to local VITS / Kokoro (offline). - Licenses of the third-party code inlined into the published package (
msedge-ttsand its dependencies) are inTHIRD_PARTY_NOTICES.md;pnpm auditis clean at release time (CIauditjob keeps watching).
Compatibility declaration (what the plugin market's version filter reads)
package.json declares engines.dsh = ">=0.1.1-rc.2" (the market uses it for "filter by host version"). Verified range: 0.1.1-rc.2 → 0.2.1-alpha.1 (including an Electron host equivalent to the official desktop app). The upper bound is open; a weekly CI check watches upstream releases and opens an issue for any uncovered version. peerDependencies only lists @deepseek-ai/cordis and react (provided by the host).
🚧 Known limitations
- Barge-in relies on browser echo cancellation (
echoCancellation); loud speaker volume may leak into the mic (no JS-level AEC) Ctrl+Shift+Voverrides the browser's "paste as plain text" shortcut (normalCtrl+Vpaste still works)- The recognition model prioritizes Simplified Chinese; recognition quality is affected by ambient noise
- Browser autoplay policy: reading requires prior user interaction on the page (clicking the mic satisfies it); if the browser blocks playback and the status bar shows no hint, make sure the page is foregrounded and not muted
- The wake word is a lightweight implementation (text matching on the streaming transcript, not a dedicated KWS engine): it may lag or misfire in noisy environments; the wake word itself never enters the chat (the buffer is dropped on hit)
- In hold mode, switching windows/tabs while holding abandons the segment (prevents continuous recording); come back and hold again
- The hero (new-session empty state) has no voice entry: voice mode is a session-level feature; enter a session first and use the mic button in the input toolbar
- The preview request timeout uses
AbortSignal.timeout(Chrome 103+ / Firefox 100+ / Safari 16+); on older browsers clicking preview immediately shows a failure hint — an expected degradation
🛠️ Troubleshooting
| Symptom | Fix |
|---|---|
| Mic click does nothing, red hint in the status bar | The browser denied mic permission: allow it in the address bar and retry |
Status bar stuck on Loading model… x% |
Check the network; the model is large (160 MB) — npm run prefetch first; on mainland networks set modelHost to https://hf-mirror.com |
Status bar shows Model download failed (<file>) |
Both mirrors are unreachable: check network/proxy and re-enter voice mode (resumable) |
| Local Kokoro on the official desktop app shows "failed" and asks for Node.js | The desktop host is Electron, where Kokoro's native add-on cannot run: install Node.js >= 18 on PATH (or set DSHVM_NODE to its path) and retry, or use the default Edge / local VITS (unaffected on desktop) |
| Caption appears (overlay) but no sound | Check system volume/output; if autoplay is blocked, click anywhere on the page and retry |
Status bar shows 朗读连接失败:正在重试… |
Edge TTS unreachable (overseas service), auto-retries; if it persists, check network/proxy |
| Poor recognition | Get closer to the mic, reduce ambient noise; if echo remains, raise the interrupt sensitivity by one step |
| Hold mode has no effect | Make sure hold mode is active and you're in voice mode (button shows 按住说话); the browser window must be in the foreground |
| Preview button reports synthesis failure | Edge TTS unreachable (overseas) or the ShortName doesn't exist: verify the name (node scripts/list-voices.mjs lists all) and retry later |
| Caption is hidden behind the input box | Default captionMaxWidth=1 (70vw) + captionFontSize=0 (12px) can overlap the bottom input on narrow viewports; raise the tier or click the caption's × to dismiss |
Yield behavior is wrong (saying 嗯 doesn't pause / real speech gets hard-barge) |
Short backchannel words (嗯 / 对) auto-pause 1.5 s then reading resumes; continuing to speak triggers hard barge-in; disable backchannelYield to restore pre-change behavior (ADR-0008) |
Says 嗯 but no yield happens |
Confirm backchannelYield=true (default on); in hold mode, continuing to talk within 1.5 s of release triggers hard barge-in |
| Auto-exits after 5 minutes idle (don't want) | Raise idleTimeoutMinutes (default 5 min, reading counts as activity) |
🛣️ Roadmap
Full backlog (43 P0-P3 items) at docs/competitive/backlog.md.
- ✅ Done (v0.7.7): 11 batches of comprehensive fixes (caption tiers / yield semantics / model prewarm / defaults micro-adjust / dead-code cleanup / etc.)
- 🚧 P0 (near-term): ADR-0003 client-side VAD / ADR-0006 first-level probe wired to manual / F1 emotion DSL full roll-out
- 📋 P1 (mid-term): MCP
voice_*toolset / card form draft validate / status-bar idle polish - 💡 P2 (far-term): Voice cloning (user-deferred) / ADR-0004 WebSocket transport
- ⏸️ Deferred: xAI fallback / C1 persona layer (user-deferred)
🛠️ Development
Dependency discipline (important)
Three categories, each with its own home:
- Third-party runtime deps (
msedge-tts/sherpa-onnx) and registry-resolvable framework packages (@deepseek-ai/schemastery) →dependencies. schemastery is a public npm package and the dsh host platform does not shadow it internally, so installing it into a profile causes no version conflicts. - Host framework packages (
@deepseek-ai/cordis/@deepseek-ai/dsh-web/react) →peerDependencies. Host packages are provided by the dsh runtime; putting them in dependencies makes dshmarket treat it as "shadowing host versions" and blocks marketplace upgrades. Peer versions must match the current dsh runtime (locally: cordis^4.0.1, dsh-web^0.1.0-rc.6 || ^0.1.1-rc.0, react^18.2.0), and be bumped when dsh is upgraded. - Type-only references (
@deepseek-ai/dsh-settings/dsh-host-webserver/dsh-llm) → no runtime imports between instances (import type+ esbuild stripping), no declaration needed; during development the types link via pnpmfile:to the local dsh distribution's node_modules (the registry's rc.1 type snapshot lags behind the distribution; the distribution types are the runtime truth).
dependencies must contain only true third-party runtime deps — never host-shared packages; after
changing deps, run npm pack --dry-run and pnpm test as regression.
Build & test
pnpm install && pnpm build # esbuild: lib/index.js (host) + lib/client.js (browser)
pnpm test # segmenter/wakeword unit tests + pre-release self-check (no network)
node test/hold-e2e.js # hold-mode acceptance (standalone browser, /asr route interception)
systemctl restart dsh # Linux; restart the dsh process on other platforms
Note: dsh installs the plugin as a pnpm
file:link (directory copy); afternode build.mjsyou must copylib/client.jsto<profile>/node_modules/dsh-voice-mode/lib/and restart dsh before the browser picks up the new bundle.
Structure
src/index.ts host: single-active pointer, llm/stream tap, SSE, settings registration
src/asr-host.ts host: zipformer2 streaming ASR + lazy model download (.part resume)
src/tts-queue.ts host: per-session TTS queue + epoch barge-in
src/segmenter.ts host: sentence segmentation (markdown stripping + terminating punctuation)
src/client.tsx client: mic button + status bar + reading overlay + barge-in
src/asr.ts client: getUserMedia + RMS VAD + partial polling
scripts/prefetch.mjs model pre-download (cross-platform cache dir + resume)
test/segmenter.test.mjs sentence segmentation unit tests
test/wakeword.test.mjs wake-word matching unit tests
test/verify-client.mjs pre-release self-check (bundle manifest/exports/shape)
test/hold-e2e.js hold-mode end-to-end acceptance (standalone browser)
scripts/list-voices.mjs print all Edge TTS voices (source of the voice table)
Integration probes (hold-e2e.js, spoken-prompt-rpc.sh, spoken-toggle-ui-check.js) live in the repo root test/, outside this npm package.
📄 License
Links
More in this category
PolinniZhong/dsh-omi-voice★ 74
In-chat read-aloud for DeepSeek Harness: tap to read, pause and resume AI replies with natural Doubao TTS voices (BYOK), reading only the final answer with code, tables and diagrams filtered; local engine, plugin keeps no API key.
PensiveFei/dsh-voice-scribe★ 35
Voice input for the web UI: tap Alt (or Alt+Space) to start/stop dictation, browser Web Speech (zero-config) or OpenAI-compatible cloud ASR, optional polish through DSH-configured LLM, settings UI.
1624318455/dsh-plugin-tts★ 24
Reads assistant replies aloud via free Edge TTS or your own RVC voice models: read-aloud buttons + auto-read, adaptive chunked progressive playback (gapless long reads), one-click voice-pack installs from a registry, and a portable RVC runtime.
WizisCool/dsh-ears★ 22
Voice input plugin for DeepSeek Harness (dsh): a microphone button in the composer turns speech into a draft transcript, with a choice of speech-recognition backends, optional polish through dsh own LLM routes, and a native settings page.
PerryLink/dsh-talk★ 17
Voice I/O for DeepSeek Harness — speech-to-text and text-to-speech over the microphone and audio output.
ppy-web/dsh-plugin-xiaomi-mimo-tts★ 15
Adds Xiaomi MiMo text-to-speech to DSH Web with assistant-message read-aloud, PCM streaming, preset and custom voice design, browser speech fallback, playback controls, and optional UI sounds.
Community comments
Comments are public GitHub Discussions. Loading them connects to GitHub and Giscus; a GitHub account is required to post.