Voice input plugin for DeepSeek Harness (dsh): a microphone button in the composer turns speech into a draft transcript, with a choice of speech-recognition backends, optional polish through dsh own LLM routes, and a native settings page.
Install
# from npm (prebuilt)
dsh plugin --profile web add dsh-ears
# from GitHub (first run asks for allowBuilds approval — follow the hint, retry)
dsh plugin --profile web add github:WizisCool/dsh-ears
Any plugin you install runs third-party code with your own permissions — it can read your files, use your credentials, and reach the network, and tool approvals don’t sandbox it. GitHub-sourced plugins also run build scripts at install time — pnpm blocks those until you allow them, so an install can stop with ERR_PNPM_GIT_DEP_PREPARE_NOT_ALLOWED or ERR_PNPM_IGNORED_BUILDS; dsh prints the exact key to add under allowBuilds in your profile’s pnpm-workspace.yaml, and the install works on the next run. Allowing a build is a trust decision: only install sources you trust, and pin a commit (github:owner/repo#sha).
README
https://github.com/user-attachments/assets/1363768e-a393-44bd-a008-1ce2055cac41
dsh-ears adds voice input and LLM-powered text polishing to DeepSeek Harness. It supports in-browser Web Speech, local Whisper, and popular speech-to-text (ASR) APIs.
Install
Prerequisites: DeepSeek Harness >=0.2.0-rc.1, and Node.js ^22.19.0 || >=24.0.0.
Install from npm
Web UI:
dsh plugin --profile web add dsh-ears
Desktop application: install from its plugin manager. The desktop app owns a reserved profile named desktop and the dsh CLI refuses to manage it, so there is no command-line equivalent. dsh-ears needs no separate desktop build — both surfaces load the same web client bundle.
Still using dsh 0.1.x? dsh-ears
0.4.1requires dsh>=0.2.0-rc.1, and dsh refuses to install it on an older host. Stay on the matching earlier line:# dsh 0.1.2 through 0.1.6 dsh plugin --profile web add "dsh-ears@<0.4.0" # dsh 0.1.7 (dsh-ears 0.4.0 requires dsh 0.1.7-rc.2 or later) dsh plugin --profile web add "dsh-ears@<0.4.1" # dsh 0.1.1 dsh plugin --profile web add "dsh-ears@<0.3.0"
Install from source
git clone https://github.com/WizisCool/dsh-ears.git
cd dsh-ears
pnpm use:platform
pnpm install
pnpm build
dsh plugin --profile web add "$PWD"
# Windows cmd: use "%CD%"; PowerShell expands $PWD directly
pnpm use:platform switches node_modules to the native dependency tree for the active platform (kept separate for Windows and Linux). Run it on a fresh clone and whenever switching between platforms.
After installation, refresh the Web UI. A microphone icon appears to the right of the composer.
Update
dsh plugin --profile web add dsh-ears
add resolves the latest version from npm, so the same command works from any installed version. After updating, restart dsh web to load the new server code and refresh the Web UI; you can also check for new versions in the "About" panel on the settings page.
Uninstall
dsh plugin --profile web remove dsh-ears
Refresh the Web UI after uninstalling.
Usage
- Click the microphone icon or press
Ctrl+Shift+Space(configurable in settings). - Start speaking.
- Press the shortcut again or click the microphone to stop recording and start transcription.
- With polishing enabled, the raw transcript appears in the draft first. Once polishing completes, it replaces the text while preserving any manual edits made in the meantime.
- Review the text and send manually.
Transcription results are always written to an editable draft and are never sent automatically. If the selected backend is not ready, the microphone icon is disabled — hover over it to see why.
Recognition backends
| Backend | How it works | Requirements |
|---|---|---|
| Web Speech | In-browser real-time recognition | Default backend; Chromium-based browser |
| Local Whisper | Local transcription after recording stops | Download GGML models in settings |
| Groq | Groq Whisper API | API key |
| Deepgram | Pre-recorded audio or live audio streaming | API key, model name (e.g. nova-3) |
| Alibaba Cloud Model Studio | DashScope synchronous transcription | API key, model name; max 300 s per recording |
| Tencent Cloud | Recording file recognition or real-time WebSocket | AppID, SecretID, SecretKey, engine_type |
| Xiaomi MiMo | Speech Recognition API, supporting standard API or Token Plan | API key, model name (e.g. mimo-v2.5-asr); Token Plan requires a regional cluster |
| SiliconFlow (CN) | Audio Transcription API | API key, model name (e.g. FunAudioLLM/SenseVoiceSmall) |
| Volcengine | Doubao audio file recognition or Doubao one-way streaming ASR | API key (new-console X-Api-Key), resource ID |
| Custom OpenAI-compatible | Sends to a specified /audio/transcriptions endpoint |
Endpoint URL, API key, model name |
Local Whisper: Runs whisper.cpp locally via
@fugood/whisper.nodewithout Python, FFmpeg, or external dependencies. Supports downloading and managing GGML models (tiny(default) throughturbo) in settings, withdefault(auto),vulkan, orcudaacceleration.Volcengine: Accepts only the new-console API Key (
X-Api-Key) authentication (legacy AppID + Access Token is not supported). Obtain keys from the console API Key page.
Polishing
Enabled by default. Leave both provider and model empty to use dsh's default Agent model (including its default reasoning settings); you can also select any model configured in dsh. LLM credentials reuse dsh's existing configuration. Reasoning effort is configurable.
The default prompt removes filler words, fixes common ASR errors, handles self-corrections, and formats lists. Customize the prompt or view the default content in settings. If polishing fails or is cancelled, the raw transcript is preserved.
Settings
After installation, a dsh-ears section appears in the Web UI settings page, organized into four tabs:
| Tab | Configurable options |
|---|---|
| General | Settings display name, voice shortcut toggle and key combination, audio cues toggle, max recording duration (1–600 s, default 120 s) |
| Recognition | Backend selection, language, local Whisper model and acceleration, cloud providers and credentials |
| Polishing | Toggle, LLM provider and model, reasoning effort, custom prompt (up to 4,000 characters) |
| About | Version, license, dsh compatibility range, update checker |
Known limitations
- The Web Speech backend relies on the browser implementation; audio may be sent to browser vendor servers for processing and is not a strictly local/offline solution.
- Alibaba Cloud Model Studio (Bailian) recordings are capped at 300 seconds per audio.
- Deepgram Flux models require the Listen V2 protocol and are currently unsupported.
- Local Whisper has a single-recording limit of 24 MB and a 120-second transcription timeout.
- Acceleration backend (
default/vulkan/cuda) is locked after the first native module load; switching requires restartingdsh web.
Local development
pnpm use:platform
pnpm install
dsh plugin --profile web add "$PWD"
# Windows cmd: use "%CD%"; PowerShell expands $PWD directly
pnpm check # type-check
pnpm test # run tests
pnpm build # build
pnpm dev:config # build and write the HMR config
pnpm dev:web # start dsh web
Run pnpm dev:watch in a second terminal while developing for live rebuilds.
After modifying frontend UI code, simply refresh the browser. When modifying server-side code, settings registration, Remote descriptors, or schemas, restart dsh web and then refresh.
Docs
License
Community Links
- LINUX DO — A new ideal community
Star History
Links
More in this category
PolinniZhong/dsh-omi-voice★ 74
In-chat read-aloud for DeepSeek Harness: tap to read, pause and resume AI replies with natural Doubao TTS voices (BYOK), reading only the final answer with code, tables and diagrams filtered; local engine, plugin keeps no API key.
PensiveFei/dsh-voice-scribe★ 34
Voice input for the web UI: tap Alt (or Alt+Space) to start/stop dictation, browser Web Speech (zero-config) or OpenAI-compatible cloud ASR, optional polish through DSH-configured LLM, settings UI.
1624318455/dsh-plugin-tts★ 21
Reads assistant replies aloud via free Edge TTS or your own RVC voice models: read-aloud buttons + auto-read, adaptive chunked progressive playback (gapless long reads), one-click voice-pack installs from a registry, and a portable RVC runtime.
PerryLink/dsh-talk★ 16
Voice I/O for DeepSeek Harness — speech-to-text and text-to-speech over the microphone and audio output.
qishuilalala/dsh-voice-mode#dsh-voice-mode★ 15
Full-duplex voice mode for the DeepSeek Harness Web UI: toggle (2s-pause auto-send) or hold-to-talk dictation into an editable draft with zipformer2 streaming ASR, optional wake word; sentence-by-sentence Edge TTS read-aloud with live captions, and speaking interrupts playback and the running turn (true barge-in). On-device ASR, no API key.
beiyege-01/dsh-voice-ai-girlfriend-plugin★ 12
Voice AI girlfriend for the Web UI: FunASR mic input, Qwen3-TTS spoken replies, companion animation window, and two-way QQ chat (text/voice/image push) via NapCat.
Community comments
Comments are public GitHub Discussions. Loading them connects to GitHub and Giscus; a GitHub account is required to post.