Voice input for the web UI: dictation chunked by pauses and voice messages, each with its own provider fallback chain (Deepgram, Groq, HuggingFace, local whisper.cpp, or an OpenAI-compatible endpoint).
Install
# from npm (prebuilt)
dsh plugin --profile web add @goodandready/dsh-voice
# from GitHub (first run asks for allowBuilds approval — follow the hint, retry)
dsh plugin --profile web add github:GooDAnDReaDY/dsh-voice
Any plugin you install runs third-party code with your own permissions — it can read your files, use your credentials, and reach the network, and tool approvals don’t sandbox it. GitHub-sourced plugins also run build scripts at install time — pnpm blocks those until you allow them, so an install can stop with ERR_PNPM_GIT_DEP_PREPARE_NOT_ALLOWED or ERR_PNPM_IGNORED_BUILDS; dsh prints the exact key to add under allowBuilds in your profile’s pnpm-workspace.yaml, and the install works on the next run. Allowing a build is a trust decision: only install sources you trust, and pin a commit (github:owner/repo#sha).
README
Requirements
Browser Microphone & Secure Context\nModern browsers require a Secure Context (https:// or http://localhost / http://127.0.0.1) for microphone capture (navigator.mediaDevices.getUserMedia). When accessing DSH over a local network (e.g., http://192.168.1.x:3080 via dsh-lanmode), browsers will block microphone access by default. Serve DSH behind an HTTPS reverse proxy or configure browser security flags (chrome://flags/#unsafely-treat-insecure-origin-as-secure) for LAN IP access.\n
⚡ Overview
dsh-voice brings voice superpowers to the DeepSeek Harness Web UI. Whether you need hands-free real-time streaming dictation segmented on natural breath pauses or crisp voice notes with keyboard/mouse Push-to-Talk gestures, dsh-voice ensures your audio is never lost thanks to automatic multi-provider fallback chains.
graph LR
subgraph Client [Browser Web UI]
Mic[🎙️ Dictation Mic] -->|VAD Cut on Pause| Stream[Audio Chunks]
Wave[🌊 Voice Message] -->|Hold / Release| PTT[Push-to-Talk]
end
subgraph Host [DSH Host Backend]
Stream --> FFMPEG[ffmpeg 16kHz Transcoder]
PTT --> FFMPEG
FFMPEG --> Chain{Fallback Chain}
Chain -->|1st Priority| P1[Deepgram / Nova-2]
Chain -.->|On Rate Limit / 429| P2[Groq / Whisper Turbo]
Chain -.->|On Failure| P3[Local whisper.cpp / Offline]
end
subgraph Output [Target]
P1 --> Composer[💬 Web Composer / Chat]
P2 --> Composer
P3 --> Composer
end
style Client fill:#1e1e2e,stroke:#89b4fa,stroke-width:2px,color:#cdd6f4
style Host fill:#181825,stroke:#cba6f7,stroke-width:2px,color:#cdd6f4
style Output fill:#11111b,stroke:#a6e3a1,stroke-width:2px,color:#cdd6f4
✨ Key Features
- 🎙️ Streaming Dictation with VAD: Speech is automatically sliced at natural pauses (
vadSilenceMs, default 700ms) and typed into the composer in real time. - 🌊 Voice Notes with Cancel Window: Record your thought and have it automatically dispatched to the agent after a safety countdown (
autoSendMs, default 4000ms). - 🎮 Tactile Push-to-Talk:
- Mouse: Hold the wave button — releasing sends the message; dragging pointer away discards.
- Keyboard: Hold Ctrl (or custom hotkey) for hands-free speaking; press Esc to cancel.
- ⚡ Live In-Browser Captions (
browser): Web Speech API speech-to-text with real-time floating captions as you speak (Note: standard browser recognition typically streams audio to vendor servers like Google in Chrome; for true 100% offline privacy, use localwhisperorsensevoice). - 🛡️ Ironclad Multi-Provider Fallbacks: If your primary cloud provider runs out of credits or hits a 429 rate limit, requests seamlessly fail over down the chain.
- 🧠 Context Glossary Injection: Automatically extracts code variables and identifiers from your composer draft to steer STT model accuracy on technical jargon.
- 🎵 Embedded Audio Player: Preview, scrubber, and playback of your recorded voice message directly in chat and the composer dock.
- 🔇 Hardware Noise Suppression Toggle: Configurable in settings to toggle browser-level noise suppression, echo cancellation, and auto gain control.
- 📊 Provider Latency & Health Dashboard: Live visual telemetry of provider latency (ms), success rates, and errors directly within the settings UI.
- 🔒 Zero API Key Leakage: Keys are resolved on the host via
ctx.credentials(credentialRef) and never transmitted to browser clients. - 🖥️ Offline Local Whisper Server: Automatically boots and manages whisper.cpp (
whisper-server) with on-the-flyffmpegtranscode. - ⚡ SenseVoice-ONNX / Sherpa-ONNX (0.8.11, updated 0.8.32): Ultra-fast (~50–100ms) non-autoregressive local STT engine with 1-click automatic model installation to
~/.dsh/models/sensevoice. - 🎙️ Gated Turn-Taking & Echo Prevention (0.8.32): Automatically mutes the mic while the assistant is speaking (
dsh:tts:start/stop) and supports speech-triggered Barge-In. - 👻 Live Ghost Interim Preview (0.8.32): Visual real-time preview of spoken words before final chunk transcription.
- ⏱️ SVG Silence Ring Timer (0.8.32): Circular animated countdown indicator in the composer dock during the pending message delay.
- 🗣️ Hands-free Voice Actions (0.8.32): Spoken commands ("send", "cancel", "clear", "new line") trigger UI actions directly.
- 🧩 Structured Prompt Voice Selection (0.8.32): Spoken options automatically select buttons in interactive assistant choice prompts.
- 📖 Developer Lexicon & IT Jargon Correction (0.8.32): Phonetic normalization for developer slang (GitHub, Docker, Kubernetes, pnpm) with customizable settings dictionary.
- 🌊 Liquid Wave & Dynamic Orb Visualizer (0.8.12): Smooth animated audio visualization in the recording pill with real-time mic volume reactivity. Switch between organic multi-layer liquid waves, pulsating radiant orb, classic bars, or off.
📝 What's New in 0.8.32
10-point feature epic & SenseVoice 1-Click Installer (Issue #129):
| Area | Feature | Description |
|---|---|---|
| Local ASR | 1-Click SenseVoice-Small | Automated download and extraction of model.int8.onnx and tokens.txt directly to ~/.dsh/models/sensevoice via host loopback endpoint. |
| Turn-Taking | Gated Mode | Microphone input is muted when assistant audio starts (dsh:tts:start) and unmuted on dsh:tts:stop, eliminating acoustic feedback. |
| Interruption | Barge-In | Speaking immediately signals assistant TTS to pause and abort current speech playback. |
| Feedback | Ghost Interim Preview | Semi-transparent live text preview in composer pill during recognition before sentence finalizing. |
| Countdown | SVG Silence Ring | Circular SVG progress ring visually ticking down pending auto-send delay. |
| Control | Hands-free Actions | Spoken control words ("send", "cancel", "clear", "new line") trigger actions instead of becoming message text. |
| Prompts | Structured Prompts | Spoken replies automatically match and submit options in active harness interactive prompts. |
| Vocabulary | IT Jargon Normalizer | Auto-corrects spoken developer slang to canonical spelling (GitHub, Docker, Kubernetes, etc.) with custom dictionary in Settings. |
| Configuration | Individual Toggles | Dedicated switches for every enhancement in Settings card with native --dsw-alias-* token styling. |
📝 What's New in 0.8.19
Quality batch after the v0.8.18 review (Gitea #79–#87, PR #88):
| Area | Change |
|---|---|
| Settings placeholders | Model/path hints resolve through locale at render time — no frozen i18n keys (#79) |
| Model field hints | whisperModel / sensevoiceModel show path-to-model copy, not binary-autostart hints (#82) |
| Visualizer | Liquid wave / dynamic orb read DSH theme tokens instead of hardcoded hex (#80) |
| Style isolation | Injected CSS uses data-dsh-plugin="dsh-voice" so neighbour HMR cleanup cannot strip styles (#81) |
| Source language | Code comments, errors, and tests are English; EN edit-command phrases added. Spoken RU STT patterns remain for recognition (#86) |
| Locales | Changed: the full inline ru dictionary was removed. English is the only bundled locale. Install the translation plugin for Russian UI (#85) |
| Client source | Browser client is built from ordered lib/client-src/*.js fragments via npm run build:client (#87) |
| Process docs | Added docs/design/DESIGN.md, index.md, project AGENTS.md, docs/testing/unit.md (#83) |
[!IMPORTANT] v0.8.19 locale behavior: without a translation plugin the Web UI stays in English. Session/edit spoken command phrases still match Russian speech for STT where documented; interface labels do not ship a second dictionary.
🎮 Four Ways to Speak
| Mode | Gesture / Trigger | Behavior |
|---|---|---|
| Dictation | Click 🎙️ Mic | Speech is sliced on pauses (vadSilenceMs) and typed live into composer |
| Voice Message | Click 🌊 Wave | Records until stopped, then sends after cancel window (autoSendMs) |
| Mouse PTT | Hold 🌊 Wave | Records while held; release sends message, drag off button to discard |
| Keyboard PTT | Hold Ctrl | Hands-free recording; release sends message, press Esc to discard |
[!TIP] You can customize the keyboard modifier in settings (
hotkey:Control,Alt,Shift, or anyKeyboardEvent.code).
🛠️ Supported Providers Matrix
| Provider Key | Service Backend | Default Model | Credential Ref | Features & Notes |
|---|---|---|---|---|
browser-webgpu |
In-Browser WebGPU/WASM Whisper | onnx-community/whisper-tiny |
None | 100% private offline client-side GPU transcription; falls back to server chain |
browser |
Web Speech API | Native Browser | None | Zero latency, floating live captions in Chrome (vendor server-based; not offline) |
deepgram |
Deepgram API | nova-2 |
DEEPGRAM_API_KEY |
Ultra-fast cloud transcription |
groq |
Groq Whisper | whisper-large-v3-turbo |
GROQ_API_KEY |
Near-instant inference speed |
hf |
HuggingFace Inference | openai/whisper-large-v3 |
HF_TOKEN |
High-accuracy open Whisper |
local-whisper |
Local whisper.cpp | Server defined | None | 100% private, offline, no internet needed |
sensevoice |
SenseVoice-ONNX / Sherpa-ONNX | SenseVoiceSmall |
None | Ultra-fast (~50ms) local STT; supports Rockchip RK3588 NPU (rknpu), OpenVINO, CUDA |
🚀 Ready-Made Presets (Plug & Play)
Just specify the name in your fallback chain and add the corresponding API key:
openai(whisper-1) →OPENAI_API_KEYsiliconflow(SenseVoiceSmall) →SILICONFLOW_API_KEYmistral(voxtral-mini-latest) →MISTRAL_API_KEYopenrouter(google/gemini-2.5-flash) →OPENROUTER_API_KEYdeepinfra(whisper-large-v3-turbo) →DEEPINFRA_API_KEYfireworks(whisper-v3-turbo) →FIREWORKS_API_KEY
📦 Quick Installation
dsh plugin --profile web add @goodandready/dsh-voice
[!IMPORTANT] Restart DSH Web UI after installation (
systemctl --user restart dsh-web) and refresh your browser tab.
⚙️ Configuration
Open Settings → Plugins → Plugin settings → Voice in the Web UI:
- id: dsh-voice
config:
dictation:
language: "" # auto-detect by default; or specify "en", "ru", "zh", etc.
vadSilenceMs: 700
chain:
- provider: deepgram
- provider: groq
- provider: local-whisper
message:
language: "" # auto-detect by default; or specify "en", "ru", "zh", etc.
autoSendMs: 4000
chain:
- provider: openai
- provider: local-whisper
hotkey: Control
autoStart: true
whisperModel: /models/ggml-medium-q8_0.bin
Note on language handling: Setting
language: ""(default) allows providers to auto-detect the spoken language. Spoken numeral normalization (wordsToDigits) and standard Russian tech slang correction (DEFAULT_JARGON_DICTIONARY) activate conditionally only when Russian is selected or auto-detected in the transcript, ensuring non-Russian speech is never distorted.
🤖 Agent Tool & HTTP API
Agent Tool (transcribe_audio)
Registers transcribe_audio(file_path, language?) in ctx.tools, allowing agents to analyze audio files, interview recordings, and voice notes directly from disk. Path containment is strictly enforced against allowed directories (allowedAudioDirs, ~/.dsh, tmpdir, cwd) with symlink traversal escape prevention and magic-byte audio format validation.
Internal HTTP Endpoints
POST /dsh-voice/transcribe—{ dataBase64, mimeType, mode }→{ ok, text, provider, tookMs }POST /dsh-voice/polish—{ text }→{ ok, text }GET /dsh-voice/status— Returns daemon status, active chains, SenseVoice andeffectiveSensevoiceProvider.GET /dsh-voice/config— Returns live plugin configuration snapshot.PUT /dsh-voice/config— Updates and persists plugin configuration across network.GET/POST /dsh-voice/sensevoice-installer— SenseVoice 1-click model installation status and trigger (protected byisTrustedCaller, model paths redacted).POST /api/dsh-voice/update— Triggers plugin self-update to latest compatible npm version (protected by local caller verification and profile lock check).GET /api/dsh-voice/update— Returns update check status (current version, latest version, updateAvailable).
🔒 Profile Lock & Controlled Operator Recovery Policy
When updating or installing plugins, DeepSeek Harness profiles coordinate concurrent operations using <profile-dir>/package.json.lock (for example, ~/.dsh/profiles/<profile>/package.json.lock).
- Canonical Contender Policy: The contender process never deletes or unlinks an existing lockfile. Automatic stale lock removal by contenders is prohibited to eliminate TOCTOU (Time-of-Check to Time-of-Use) race conditions and prevent unintended lockfile clobbering.
- Diagnostics: If
<profile-dir>/package.json.lockis present, update attempts return409 Conflictwith clear diagnostics:- If held by an active process, the PID is reported.
- If filesystem inspection encounters access errors (
EACCES,EIO), the operation fails closed with the error code preserved. Filesystem errors require inspecting filesystem health, permissions, or mount options — do not remove the lockfile in response toEACCES/EIO.
- Controlled Operator Recovery Workflow:
If a previous installation was forcefully interrupted (e.g. system crash, OOM kill) leaving a stale lockfile, recovery is strictly an operator action and must follow this controlled procedure:
- Establish Exclusive Maintenance: Ensure no concurrent processes, CLI commands (
dsh plugin add), web UI operations, or automated updaters are running or scheduled to start for the target profile. - Verify Quiescence and Ownership: Check for active processes (
pgrep,ps,fuser, orlsof) to confirm that no live process is currently performing package operations in the target profile directory. Lock age alone or observing an unparsed PID is not sufficient justification to delete a lock while another installer could start. If quiescence or non-ownership cannot be verified, do not remove the lock. - Remove Confirmed Orphan Lock: Only after confirming the lock is an orphan and exclusive maintenance is established, remove the specific profile lock without recursive or force flags:
rm ~/.dsh/profiles/<profile>/package.json.lock - Resume Operations: Re-enable installations and retry the plugin update.
- Establish Exclusive Maintenance: Ensure no concurrent processes, CLI commands (
📄 License
MIT © GooDAnDReaDY
Links
More in this category
PolinniZhong/dsh-omi-voice★ 75
In-chat read-aloud for DeepSeek Harness: tap to read, pause and resume AI replies with natural Doubao TTS voices (BYOK), reading only the final answer with code, tables and diagrams filtered; local engine, plugin keeps no API key.
PensiveFei/dsh-voice-scribe★ 35
Voice input for the web UI: tap Alt (or Alt+Space) to start/stop dictation, browser Web Speech (zero-config) or OpenAI-compatible cloud ASR, optional polish through DSH-configured LLM, settings UI.
1624318455/dsh-plugin-tts★ 24
Reads assistant replies aloud via free Edge TTS or your own RVC voice models: read-aloud buttons + auto-read, adaptive chunked progressive playback (gapless long reads), one-click voice-pack installs from a registry, and a portable RVC runtime.
WizisCool/dsh-ears★ 22
Voice input plugin for DeepSeek Harness (dsh): a microphone button in the composer turns speech into a draft transcript, with a choice of speech-recognition backends, optional polish through dsh own LLM routes, and a native settings page.
PerryLink/dsh-talk★ 17
Voice I/O for DeepSeek Harness — speech-to-text and text-to-speech over the microphone and audio output.
ppy-web/dsh-plugin-xiaomi-mimo-tts★ 17
Adds Xiaomi MiMo text-to-speech to DSH Web with assistant-message read-aloud, PCM streaming, preset and custom voice design, browser speech fallback, playback controls, and optional UI sounds.
Community comments
Comments are public GitHub Discussions. Loading them connects to GitHub and Giscus; a GitHub account is required to post.