Voice input for DeepSeek Harness: speak into the microphone and the recognized text is submitted as a normal chat message, via local or browser speech recognition.
Install
# from GitHub (first run asks for allowBuilds approval — follow the hint, retry)
dsh plugin --profile web add github:3274375092/dsh-voice
Any plugin you install runs third-party code with your own permissions — it can read your files, use your credentials, and reach the network, and tool approvals don’t sandbox it. GitHub-sourced plugins also run build scripts at install time — pnpm blocks those until you allow them, so an install can stop with ERR_PNPM_GIT_DEP_PREPARE_NOT_ALLOWED or ERR_PNPM_IGNORED_BUILDS; dsh prints the exact key to add under allowBuilds in your profile’s pnpm-workspace.yaml, and the install works on the next run. Allowing a build is a trust decision: only install sources you trust, and pin a commit (github:owner/repo#sha).
README
English | 中文
A voice input plugin for DeepSeek Harness: click 🎤 in the web UI (or press a hotkey), speak, and the recognized text is submitted as a normal chat message. Input only — it never touches the agent preset/persona, so it behaves like "another input method" in every mode.
Features
- 🎤 Voice input: microphone button (platform design-system UI) + configurable global hotkey (Ctrl+Space by default)
- ⚡ Live recognition: streaming partial transcripts are echoed while you speak; VAD finalization commits on stop (0.6s tail padding keeps sentence endings)
- 🧠 Adaptive dual engine: host-native ASR (sherpa-onnx-node zipformer2 + silero VAD, offline/private) with automatic fallback to browser Web Speech (zero extra dependencies)
- 🔌 Preset-agnostic: does not touch the persona/system prompt; works with code/standard/minimal/custom presets
- 📦 Optional models: zero-config out of the box; run
dsh-voice-modelswhen you want native offline recognition
Installation
# Plugin
dsh plugin --profile web add @nn12138/dsh-voice
# Optional: offline native recognition (the plugin does not auto-install this runtime)
dsh plugin --profile web add sherpa-onnx-node
dsh-voice-models # one-shot model download (~100MB) → ./dsh-voice-models
# Optional: configuration (edit ~/.dsh/profiles/web/cordis.patch.yml)
- id: voice
config:
modelDir: './dsh-voice-models' # native ASR model directory
hotkey: 'ctrl+space' # global hotkey
vadThreshold: 0.3 # lower = less clipping at sentence boundaries
tailPadSeconds: 0.6 # tail-padding duration
engine: auto # auto (default) | native | browser
The row-level config is received by the host half. engine and hotkey are synced to the browser half over the /voice.config loopback RPC, so there is no separate client config to write. auto probes host native capability: with a model it uses native; without one it falls back to Web Speech, so zero-config users keep working. Restart dsh web after changing the config.
dsh web # 🎤 button appears on the left of the composer, or press Ctrl+Space
See USAGE.md and INSTALL.md (Chinese) for details.
How it works
Browser captures mic audio (auto-resampled to 16 kHz)
→ PCM base64 chunks (256 ms) → /voice RPC channel (loopback)
→ host: silero VAD + zipformer2 streaming decode
→ partials returned per chunk (live echo) / finals committed
(VAD segmentation + 0.6s tail padding)
→ conversation service submits the text (same path as typing)
Engine selection: the host resolves the effective engine (config + model-load result) and the client consumes it via /voice.ping — native unavailable falls back to browser Web Speech. /voice.config carries the row-level engine/hotkey from host to client.
Development
pnpm install --ignore-workspace # standalone deps (no DSH monorepo needed); no install-time scripts
pnpm --ignore-workspace test # unit tests (including real-model smoke tests)
pnpm --ignore-workspace typecheck # type check
pnpm --ignore-workspace build # build (tsc host half + tsdown client half)
pnpm --ignore-workspace verify:package # pack-level checks: scripts-free install, complete runtime files
The build only ever runs on the publisher's side (prepack — at npm pack/npm publish time) and in CI before release; consumers installing @nn12138/dsh-voice from the registry never execute lifecycle scripts, so --ignore-scripts installs are complete and usable (see issue #2).
Real-model smoke tests look for the local voxelf assets and skip when absent; override with:
DSH_VOICE_MODEL_DIR (model directory) / DSH_VOICE_TEST_WAV (test wav) / DSH_VOICE_DOWNLOADED_MODELS (downloaded model directory).
Layout: src/index.ts (host half) / src/client/ (browser half) / src/core/ (recognition core) / tools/ (wire-protocol smoke tools + model downloader).
License
MIT
Links
More in this category
PolinniZhong/dsh-omi-voice★ 74
In-chat read-aloud for DeepSeek Harness: tap to read, pause and resume AI replies with natural Doubao TTS voices (BYOK), reading only the final answer with code, tables and diagrams filtered; local engine, plugin keeps no API key.
PensiveFei/dsh-voice-scribe★ 34
Voice input for the web UI: tap Alt (or Alt+Space) to start/stop dictation, browser Web Speech (zero-config) or OpenAI-compatible cloud ASR, optional polish through DSH-configured LLM, settings UI.
1624318455/dsh-plugin-tts★ 21
Reads assistant replies aloud via free Edge TTS or your own RVC voice models: read-aloud buttons + auto-read, adaptive chunked progressive playback (gapless long reads), one-click voice-pack installs from a registry, and a portable RVC runtime.
WizisCool/dsh-ears★ 21
Voice input plugin for DeepSeek Harness (dsh): a microphone button in the composer turns speech into a draft transcript, with a choice of speech-recognition backends, optional polish through dsh own LLM routes, and a native settings page.
PerryLink/dsh-talk★ 15
Voice I/O for DeepSeek Harness — speech-to-text and text-to-speech over the microphone and audio output.
qishuilalala/dsh-voice-mode#dsh-voice-mode★ 15
Full-duplex voice mode for the DeepSeek Harness Web UI: toggle (2s-pause auto-send) or hold-to-talk dictation into an editable draft with zipformer2 streaming ASR, optional wake word; sentence-by-sentence Edge TTS read-aloud with live captions, and speaking interrupts playback and the running turn (true barge-in). On-device ASR, no API key.
Community comments
Comments are public GitHub Discussions. Loading them connects to GitHub and Giscus; a GitHub account is required to post.