Mic button in the composer tool row: Web Speech API speech-to-text (Chrome/Edge), language switching, and optional auto-send, zero dependencies.
Install
# from GitHub (first run asks for allowBuilds approval — follow the hint, retry)
dsh plugin --profile web add github:0nt-one/dsh-voice-input
Any plugin you install runs third-party code with your own permissions — it can read your files, use your credentials, and reach the network, and tool approvals don’t sandbox it. GitHub-sourced plugins also run build scripts at install time — pnpm blocks those until you allow them, so an install can stop with ERR_PNPM_GIT_DEP_PREPARE_NOT_ALLOWED or ERR_PNPM_IGNORED_BUILDS; dsh prints the exact key to add under allowBuilds in your profile’s pnpm-workspace.yaml, and the install works on the next run. Allowing a build is a trust decision: only install sources you trust, and pin a commit (github:owner/repo#sha).
README
This plugin publishes its README in Chinese only.
语音输入客户端插件(client plugin)——给 DeepSeek Harness Web GUI 的聊天输入框工具行加一个麦克风按钮:点击开始说话,浏览器 Web Speech API(Chrome / Edge 内置,免费、无需 API 密钥)把语音实时转成文字填入输入框。
功能
- 麦克风按钮位于输入框工具行右侧(发送按钮左侧,
conversation.input.right插槽) - 点击开始/停止聆听;聆听时按钮红色脉冲,上方浮层实时显示转写中间结果
- 每段识别完成的最终文本自动追加到输入框草稿,可继续编辑后发送
- 浮层内可一键切换识别语言:中文 / English / 自动
- 可选自动发送:
localStorage设置dsh-voice-input-web.autoSend = "1"后,一段语音识别完成即自动发送 - 零依赖、零密钥、无服务端改动;识别完全在浏览器本地进行
与其他方案对比
| 方案 | 转写方式 | 依赖 | 成本 |
|---|---|---|---|
| 本插件(Web Speech API) | 浏览器系统语音服务 | 无 | 免费 |
| dsh-voice | agent 工具(OpenAI 兼容 ASR) | 需 ASR API key | STT 收费 |
| dsh-voice-input(SenseVoice) | 本地离线转写 | 需模型/服务 | 免费、离线 |
安装
方式一:npm 包(推荐)
dsh plugin --profile web add dsh-voice-input-web
方式二:离线安装(GitHub 源码)
将本包放入 web profile 的 node_modules:
~/.dsh/profiles/web/node_modules/dsh-voice-input-web/在
~/.dsh/profiles/web/cordis.patch.yml追加:- insert: - id: dsh-voice-input-web name: 'dsh-voice-input-web'重启
dsh --profile web,浏览器刷新页面。按钮出现在输入框工具行。
安装后插件 id / 包目录名 / localStorage key 统一使用
dsh-voice-input-web(dsh-voice-input在 npm 上已被同名 SenseVoice 方案占用)。
配置(localStorage)
| key | 值 | 默认 | 说明 |
|---|---|---|---|
dsh-voice-input-web.lang |
zh-CN / en-US / auto |
zh-CN |
识别语言 |
dsh-voice-input-web.autoSend |
"1" / "0" |
"0" |
识别完成后自动发送 |
兼容性
- 需要 Chrome / Edge(或任何支持
webkitSpeechRecognition的浏览器);Firefox 不支持 Web Speech API,会显示提示 - 页面需在
localhost/127.0.0.1或 HTTPS 下访问(浏览器安全限制) - 语音识别需要联网(由浏览器调用系统语音服务)
License
MIT
Links
More in this category
PolinniZhong/dsh-omi-voice★ 74
In-chat read-aloud for DeepSeek Harness: tap to read, pause and resume AI replies with natural Doubao TTS voices (BYOK), reading only the final answer with code, tables and diagrams filtered; local engine, plugin keeps no API key.
PensiveFei/dsh-voice-scribe★ 34
Voice input for the web UI: tap Alt (or Alt+Space) to start/stop dictation, browser Web Speech (zero-config) or OpenAI-compatible cloud ASR, optional polish through DSH-configured LLM, settings UI.
1624318455/dsh-plugin-tts★ 21
Reads assistant replies aloud via free Edge TTS or your own RVC voice models: read-aloud buttons + auto-read, adaptive chunked progressive playback (gapless long reads), one-click voice-pack installs from a registry, and a portable RVC runtime.
WizisCool/dsh-ears★ 21
Voice input plugin for DeepSeek Harness (dsh): a microphone button in the composer turns speech into a draft transcript, with a choice of speech-recognition backends, optional polish through dsh own LLM routes, and a native settings page.
PerryLink/dsh-talk★ 15
Voice I/O for DeepSeek Harness — speech-to-text and text-to-speech over the microphone and audio output.
qishuilalala/dsh-voice-mode#dsh-voice-mode★ 15
Full-duplex voice mode for the DeepSeek Harness Web UI: toggle (2s-pause auto-send) or hold-to-talk dictation into an editable draft with zipformer2 streaming ASR, optional wake word; sentence-by-sentence Edge TTS read-aloud with live captions, and speaking interrupts playback and the running turn (true barge-in). On-device ASR, no API key.
Community comments
Comments are public GitHub Discussions. Loading them connects to GitHub and Giscus; a GitHub account is required to post.