DeepSeek Harness Plugin

WizisCool/dsh-ears

Stars ★ 21 Downloads (30d) 3,032 Category Voice & Audio Added 2026-08-19 npm dsh-ears

Voice input plugin for DeepSeek Harness (dsh): a microphone button in the composer turns speech into a draft transcript, with a choice of speech-recognition backends, optional polish through dsh own LLM routes, and a native settings page.

Install

# from npm (prebuilt)

dsh plugin --profile web add dsh-ears

# from GitHub (first run asks for allowBuilds approval — follow the hint, retry)

dsh plugin --profile web add github:WizisCool/dsh-ears

Any plugin you install runs third-party code with your own permissions — it can read your files, use your credentials, and reach the network, and tool approvals don’t sandbox it. GitHub-sourced plugins also run build scripts at install time — pnpm blocks those until you allow them, so an install can stop with ERR_PNPM_GIT_DEP_PREPARE_NOT_ALLOWED or ERR_PNPM_IGNORED_BUILDS; dsh prints the exact key to add under allowBuilds in your profile’s pnpm-workspace.yaml, and the install works on the next run. Allowing a build is a trust decision: only install sources you trust, and pin a commit (github:owner/repo#sha).

README

https://github.com/user-attachments/assets/1363768e-a393-44bd-a008-1ce2055cac41


dsh-ears adds voice input and LLM-powered text polishing to DeepSeek Harness. It supports in-browser Web Speech, local Whisper, and popular speech-to-text (ASR) APIs.

Install

Prerequisites: DeepSeek Harness >=0.2.0-rc.1, and Node.js ^22.19.0 || >=24.0.0.

Install from npm

Web UI:

dsh plugin --profile web add dsh-ears

Desktop application: install from its plugin manager. The desktop app owns a reserved profile named desktop and the dsh CLI refuses to manage it, so there is no command-line equivalent. dsh-ears needs no separate desktop build — both surfaces load the same web client bundle.

Still using dsh 0.1.x? dsh-ears 0.4.1 requires dsh >=0.2.0-rc.1, and dsh refuses to install it on an older host. Stay on the matching earlier line:

# dsh 0.1.2 through 0.1.6
dsh plugin --profile web add "dsh-ears@<0.4.0"
# dsh 0.1.7 (dsh-ears 0.4.0 requires dsh 0.1.7-rc.2 or later)
dsh plugin --profile web add "dsh-ears@<0.4.1"
# dsh 0.1.1
dsh plugin --profile web add "dsh-ears@<0.3.0"

Install from source

git clone https://github.com/WizisCool/dsh-ears.git
cd dsh-ears
pnpm use:platform
pnpm install
pnpm build
dsh plugin --profile web add "$PWD"
# Windows cmd: use "%CD%"; PowerShell expands $PWD directly

pnpm use:platform switches node_modules to the native dependency tree for the active platform (kept separate for Windows and Linux). Run it on a fresh clone and whenever switching between platforms.

After installation, refresh the Web UI. A microphone icon appears to the right of the composer.

Update

dsh plugin --profile web add dsh-ears

add resolves the latest version from npm, so the same command works from any installed version. After updating, restart dsh web to load the new server code and refresh the Web UI; you can also check for new versions in the "About" panel on the settings page.

Uninstall

dsh plugin --profile web remove dsh-ears

Refresh the Web UI after uninstalling.

Usage

  1. Click the microphone icon or press Ctrl+Shift+Space (configurable in settings).
  2. Start speaking.
  3. Press the shortcut again or click the microphone to stop recording and start transcription.
  4. With polishing enabled, the raw transcript appears in the draft first. Once polishing completes, it replaces the text while preserving any manual edits made in the meantime.
  5. Review the text and send manually.

Transcription results are always written to an editable draft and are never sent automatically. If the selected backend is not ready, the microphone icon is disabled — hover over it to see why.

Recognition backends

Backend How it works Requirements
Web Speech In-browser real-time recognition Default backend; Chromium-based browser
Local Whisper Local transcription after recording stops Download GGML models in settings
Groq Groq Whisper API API key
Deepgram Pre-recorded audio or live audio streaming API key, model name (e.g. nova-3)
Alibaba Cloud Model Studio DashScope synchronous transcription API key, model name; max 300 s per recording
Tencent Cloud Recording file recognition or real-time WebSocket AppID, SecretID, SecretKey, engine_type
Xiaomi MiMo Speech Recognition API, supporting standard API or Token Plan API key, model name (e.g. mimo-v2.5-asr); Token Plan requires a regional cluster
SiliconFlow (CN) Audio Transcription API API key, model name (e.g. FunAudioLLM/SenseVoiceSmall)
Volcengine Doubao audio file recognition or Doubao one-way streaming ASR API key (new-console X-Api-Key), resource ID
Custom OpenAI-compatible Sends to a specified /audio/transcriptions endpoint Endpoint URL, API key, model name

Local Whisper: Runs whisper.cpp locally via @fugood/whisper.node without Python, FFmpeg, or external dependencies. Supports downloading and managing GGML models (tiny (default) through turbo) in settings, with default (auto), vulkan, or cuda acceleration.

Volcengine: Accepts only the new-console API Key (X-Api-Key) authentication (legacy AppID + Access Token is not supported). Obtain keys from the console API Key page.

Polishing

Enabled by default. Leave both provider and model empty to use dsh's default Agent model (including its default reasoning settings); you can also select any model configured in dsh. LLM credentials reuse dsh's existing configuration. Reasoning effort is configurable.

The default prompt removes filler words, fixes common ASR errors, handles self-corrections, and formats lists. Customize the prompt or view the default content in settings. If polishing fails or is cancelled, the raw transcript is preserved.

Settings

After installation, a dsh-ears section appears in the Web UI settings page, organized into four tabs:

Tab Configurable options
General Settings display name, voice shortcut toggle and key combination, audio cues toggle, max recording duration (1–600 s, default 120 s)
Recognition Backend selection, language, local Whisper model and acceleration, cloud providers and credentials
Polishing Toggle, LLM provider and model, reasoning effort, custom prompt (up to 4,000 characters)
About Version, license, dsh compatibility range, update checker

Known limitations

  • The Web Speech backend relies on the browser implementation; audio may be sent to browser vendor servers for processing and is not a strictly local/offline solution.
  • Alibaba Cloud Model Studio (Bailian) recordings are capped at 300 seconds per audio.
  • Deepgram Flux models require the Listen V2 protocol and are currently unsupported.
  • Local Whisper has a single-recording limit of 24 MB and a 120-second transcription timeout.
  • Acceleration backend (default / vulkan / cuda) is locked after the first native module load; switching requires restarting dsh web.

Local development

pnpm use:platform
pnpm install
dsh plugin --profile web add "$PWD"
# Windows cmd: use "%CD%"; PowerShell expands $PWD directly
pnpm check          # type-check
pnpm test           # run tests
pnpm build          # build
pnpm dev:config     # build and write the HMR config
pnpm dev:web        # start dsh web

Run pnpm dev:watch in a second terminal while developing for live rebuilds.

After modifying frontend UI code, simply refresh the browser. When modifying server-side code, settings registration, Remote descriptors, or schemas, restart dsh web and then refresh.

Docs

License

MIT

Community Links

Star History

Content from the project README on GitHub ↗

Links

More in this category

View the whole category →

Community comments

Comments are public GitHub Discussions. Loading them connects to GitHub and Giscus; a GitHub account is required to post.