输入框麦克风:点击持续监控、按住对话;浏览器语音识别逐字上屏,回复由 host Edge TTS 边生成边朗读,朗读时暂停识别防回声,点击可停止。
安装
# npm 包(预构建)
dsh plugin --profile web add @zhangbo-cn/dsh-client-ui-voice-input
# GitHub 源码(首次需按提示配置 allowBuilds 构建授权后重试)
dsh plugin --profile web add github:Zhangbo-cn/dsh-voice-input-plugin
装任何插件都等于在你的机器上跑第三方代码,权限和你本人一样大——能读你的文件、用你的凭据、访问网络,工具审批管不到它。GitHub 来源的插件还会在安装时执行构建脚本——pnpm 默认拦截,所以安装可能停在 ERR_PNPM_GIT_DEP_PREPARE_NOT_ALLOWED 或 ERR_PNPM_IGNORED_BUILDS;dsh 会打印出需要添加的确切键名,把它加进该 profile 的 pnpm-workspace.yaml 的 allowBuilds 下,重跑一次即可装上。放行构建本身就是一次信任判断:请只安装可信来源,并尽量锁定 commit(github:owner/repo#sha)。
README
该插件的 README 只有英文版本。
Composer voice control for DeepSeek Harness: a minimal linear mic button in the composer tool row that turns your speech into text — with a tap-to-monitor mode (continuous, live 逐字 streaming, send-anytime) and a hold-to-talk voice-chat mode (release to send, reply read aloud). Zero API key: recognition runs in the browser via the Web Speech API; reply reading uses the host's Edge TTS (/api/tts) with a browser speechSynthesis fallback.
dsh-plugin · TypeScript · React
Features
- Tap to monitor: click the mic, speak — text streams into the draft live (逐字输入), the mic keeps listening even in silence, and you can send or keep adding speech anytime. Tap again to stop.
- Auto-send on silence (optional): with
autoSendOnSilenceMsconfigured, tap-to-monitor becomes hands-free — once you stop talking for the configured window, the message sends itself ("tap, speak, walk away"). Continued speech cancels the pending send. - Hold to talk: press-and-hold to record a voice-chat message, release to send it; the assistant's reply is read aloud — host Edge neural TTS (
/api/tts) first, browserspeechSynthesisas fallback. - Continuous across silences: each recognition segment auto-restarts so monitoring never drops.
- Respects the composer: speech appends to the draft (base preserved); a send clears the draft cleanly without re-filling old text; monitoring continues after a send on a fresh recognizer.
- DeepSeek-blue listening state: the icon pulses in DeepSeek brand blue while listening; borderless linear icon, no clutter.
- Configurable: recognition language (default
zh-CN) and interim results.
Install
The package is a dsh.bundle installable, published on npm as @zhangbo-cn/dsh-client-ui-voice-input. One command:
dsh plugin add @zhangbo-cn/dsh-client-ui-voice-input
0.1.1+ required.
0.1.0registered the browser bundle under the wrong ModuleLoader id (@deepseek-ai/...), so Harness failed withloaded without registering "@zhangbo-cn/dsh-client-ui-voice-input". Upgrade / reinstall, then hard-refresh the Web UI.
(It also installs from the GitHub repo via dsh plugin add github:Zhangbo-cn/dsh-voice-input-plugin.)
If you develop from a DeepSeek Harness checkout, you can mount it directly in the web-app browser roster (packages/bundle/web-app/cordis.patch.yml):
- id: ui-voice-input
name: '@zhangbo-cn/dsh-client-ui-voice-input'
For reliable reply reading, also mount the host Edge TTS capability (@deepseek-ai/dsh-tts-edge), which registers /api/tts:
- id: tts-edge
name: '@deepseek-ai/dsh-tts-edge'
Without it, reply reading still works but falls back to the browser's speechSynthesis (less natural, occasionally silent on Chrome after an idle gap).
Then build the client bundle with the repo's tsdown preset:
pnpm --filter @zhangbo-cn/dsh-client-ui-voice-input run bundle
Usage
After refreshing the Web UI, the composer tool row shows a linear mic button.
Voice input (tap)
- Click the mic → the icon turns DeepSeek blue and pulses (listening).
- Speak → text appears in the input box live, word by word.
- Send anytime with the composer's send button; keep talking to add more.
- Click the mic again to stop monitoring.
- With
autoSendOnSilenceMsset, stopping speech for that long auto-sends the message — the mic stops listening, so you don't need the extra tap.
Voice chat (hold)
- Press-and-hold the mic (longer than ~250 ms) and speak.
- Release → your message is sent.
- The assistant's reply is read aloud automatically.
Reply reading after any send
A send that follows mic use (within 5 minutes) — hold or tap-monitoring + the composer send button — arms reply reading for the next assistant reply. Typed sends without recent mic use do not trigger it.
Configuration
- id: ui-voice-input
name: '@zhangbo-cn/dsh-client-ui-voice-input'
config:
language: 'zh-CN' # Web Speech recognition language tag
interimResults: true # stream live interim transcript into the draft
autoSendOnSilenceMs: 0 # auto-submit after this many ms of silence following committed speech (0 = off)
How it works
MicButton (conversation.input.left)
├─ tap → beginMonitoring()
│ → SpeechRecognition (continuous:false, interimResults) // reliable results
│ → onresult → TranscriptAccumulator → inputActions.setDraft(base + transcript)
│ → onend (silence) → auto-restart (keep monitoring) // continuous
│ → onend + committed speech + autoSendOnSilenceMs>0 → silence window → auto-submit
│ → tap again → stop
└─ hold → submitChat()
→ on release: stop + inputActions.setDraft(text) + inputActions.submit()
→ reply streams → complete sentences read aloud WHILE the model
generates (sentence-chunked queue)
→ tail (last incomplete sentence) read on finalize
→ each segment → fetch /api/tts (host Edge neural MP3)
→ play via gesture-unlocked AudioContext (else <audio> element)
→ fallback: browser speechSynthesis
- Recognition starts on pointer-down (a user gesture — required by the Web Speech API); tap vs hold is decided on release.
- The same pointer-down gesture unlocks reply audio (a shared
AudioContextis resumed), so the assistant's reply — which arrives seconds later — is exempt from the browser autoplay policy that would otherwise block a plainHTMLMediaElement.play(). - Reply reading streams: complete sentences are read aloud while the model is still generating (a sentence-chunked queue, flushed at ~30 chars for delimiter-less runs); the last incomplete sentence is read on finalize. The mic icon pulses deep blue while reading, and tapping the mic stops the reading.
- No speaker-echo: while the reply is being read, recognition is paused (the mic physically picks up the speaker), then resumes when reading finishes if monitoring was on.
continuous: falseper segment is intentional: Chrome'scontinuous: truefails to deliveronresult, so monitoring is achieved by auto-restarting segments.- The append base resets when the draft changes externally, so a send never lets stale voice text re-fill the box.
- The console logs
[dsh-voice]diagnostics for each read segment and any fallback.
Compatibility
| Browser | Mic (input, SpeechRecognition) | Reply playback (host /api/tts, fallback speechSynthesis) |
|---|---|---|
| Chrome / Edge (Windows) | ✅ Web Speech | ✅ host Edge neural MP3; browser speechSynthesis fallback |
| Safari | ✅ webkitSpeechRecognition (re-trigger on each gesture) | ✅ host Edge neural MP3 (playable); browser fallback works |
| Firefox | ⚠️ not supported — browser limitation (Mozilla has not shipped SpeechRecognition; local on-device recognition is still early-stage) |
✅ host Edge neural MP3 (playable); speechSynthesis fallback supported but less natural |
Notes:
- Firefox mic input: this is a genuine browser limitation, not a plugin issue. The plugin feature-detects and disables the mic with a "not supported in this browser" hint. A cross-browser fallback would need
MediaRecorder+ an external transcription service (out of scope for a zero-backend plugin). - Reply playback: the preferred path is the host's
/api/tts(Microsoft Edge neural voices, synthesized server-side) — reliable and natural on every browser that can play MP3. Without thetts-edgehost plugin, the client falls back tospeechSynthesis(Chrome may silently dropspeak()after an idle gap; voices are OS-default). - Mic input requires a browser with Web Speech; reply playback requires either the
tts-edgehost plugin or a browser withspeechSynthesis.
Tests
The plugin ships with two standalone-runnable suites plus a registration suite that needs the Harness monorepo.
npm install --legacy-peer-deps # peers reference monorepo-only @deepseek-ai/* packages
npm test # 36 tests: tap monitoring, auto-send on silence, hold submit,
# streaming reply reading, tap-send arming, stop-reading, markdown/emoji stripping
tests/mic-button.client.spec.tsx— component behavior (tap-to-monitor, auto-send on silence, hold-to-talk, streaming reply reading) and thesplitStreamSegments/commonPrefixLengthhelpers.tests/speech.client.spec.ts—stripMarkdownForSpeechandapplyResults.tests/apply.client.spec.ts— plugin registration; it binds real Harness services (@deepseek-ai/cordis,-runtime,-locale) that are built from the DeepSeek Harness monorepo, so it runs there (a plainvitest runfrom this package) rather than standalone. CI runs the standalone suites on Linux.
License
MIT
链接
同类插件
PolinniZhong/dsh-omi-voice★ 74
DeepSeek Harness 对话内朗读:点一下即可朗读、暂停、继续 AI 回复,豆包 TTS 自然音色(BYOK),只读最终回答并过滤代码、表格与图形,本地引擎,插件零 Key。
PensiveFei/dsh-voice-scribe★ 34
面向 Web UI 的语音输入插件:点按 Alt(或 Alt+空格)开始/停止听写,支持浏览器内置 Web Speech(零配置)或 OpenAI 兼容云端 ASR,可选经 DSH 已配置模型润色,带设置页。
1624318455/dsh-plugin-tts★ 21
用免费 Edge TTS 或你自己的 RVC 音色朗读 AI 回复:消息朗读与自动朗读、长文自适应分块渐进播放(无缝衔接)、音色包仓库一键安装、便携 RVC 运行时。
WizisCool/dsh-ears★ 21
面向 DeepSeek Harness (dsh) 的语音输入插件:输入框的麦克风按钮把语音转成草稿文本,支持多种语音识别后端,可选经 dsh 自有 LLM 路由润色,并带原生设置页。
PerryLink/dsh-talk★ 15
DeepSeek Harness 的语音输入输出:麦克风语音转文字与文字转语音。
qishuilalala/dsh-voice-mode#dsh-voice-mode★ 15
DeepSeek Harness Web UI 全双工语音对话:按钮或 Ctrl+Shift+V 进入,持续聆听(停顿自动发送)或按住说话,zipformer2 流式识别入可编辑草稿、可选唤醒词;回复按句 Edge TTS 朗读并显示实时字幕,开口即打断播放与回合(真 barge-in);识别模型本地推理、Edge TTS 在线合成,无需 API Key。
社区评论
评论公开保存在 GitHub Discussions。加载评论会连接 GitHub 和 Giscus;发表内容需要 GitHub 账号。