DeepSeek Harness Plugin

1624318455/dsh-plugin-tts

Stars ★ 21 Downloads (30d) 888 Category Voice & Audio Added 2026-08-15 npm @memef1f1y/dsh-plugin-tts

Reads assistant replies aloud via free Edge TTS or your own RVC voice models: read-aloud buttons + auto-read, adaptive chunked progressive playback (gapless long reads), one-click voice-pack installs from a registry, and a portable RVC runtime.

Install

# from npm (prebuilt)

dsh plugin --profile web add @memef1f1y/dsh-plugin-tts

# from GitHub (first run asks for allowBuilds approval — follow the hint, retry)

dsh plugin --profile web add github:1624318455/dsh-plugin-tts

Any plugin you install runs third-party code with your own permissions — it can read your files, use your credentials, and reach the network, and tool approvals don’t sandbox it. GitHub-sourced plugins also run build scripts at install time — pnpm blocks those until you allow them, so an install can stop with ERR_PNPM_GIT_DEP_PREPARE_NOT_ALLOWED or ERR_PNPM_IGNORED_BUILDS; dsh prints the exact key to add under allowBuilds in your profile’s pnpm-workspace.yaml, and the install works on the next run. Allowing a build is a trust decision: only install sources you trust, and pin a commit (github:owner/repo#sha).

README

Links


dsh-plugin-tts — Edge TTS + RVC / Index-TTS2 / CosyVoice / Cloud voices for DeepSeek Harness

A dual-sided (Host + Web UI) DeepSeek Harness plugin that reads assistant replies aloud — Microsoft Edge's free online TTS out of the box, your own RVC voice models for custom voices, local Index-TTS2 reference-audio voices, local CosyVoice zero-shot prompt clones (CosyVoice-2.0-0.5B), or paid Google Cloud TTS voices (Chinese Standard/Wavenet). Long replies stream with gapless adaptive chunked playback; voices install one-click from a voice-pack registry; a portable RVC runtime means no RVC WebUI install is needed.

📖 First time? See the user guide (执行手册) — every step covers "what / how / how to tell it worked": read-aloud, RVC voices and voice-pack downloads.

Features

  1. Read-aloud button on every finalized assistant message (in the copy / feedback / branch action row): click to speak that message (the button shows an animated equalizer), click again to stop.
  2. Auto-read toggle in the composer tool row (between the command and the access-mode buttons): when on, every newly completed assistant reply is read aloud automatically (the toggle gets a circular highlight); when off, nothing is auto-read.
  3. Voice settings panel under 设置 → 插件 → 语音:
    • TTS provider: Edge TTS (free, no API key) / custom RVC voice / local Index-TTS2 voice / local CosyVoice clone / Google Cloud TTS (paid, API key)
    • Voice: 22 live-verified Edge TTS voices (default 晓萱 zh-CN-XiaoxuanNeural); 8 Cloud voices in the cmn-CN group (default cmn-CN-Wavenet-A, free-form id allowed)
    • Sound tuning: rate / pitch / volume (0 = default, shared by Edge & Cloud)
    • RVC config: service URL + model (.pth) / index (.index) picker, base voice, voice tuning and advanced params — see the RVC Guide
    • Index-TTS2 config: service URL (default 7880) + reference-voice picker/refresh/upload + clean-text toggle + emotion control + sampling params — see the Index-TTS2 Guide
    • CosyVoice config: service URL (default 7890) + prompt-audio picker/refresh/upload + prompt transcript (required, in order and in full) + speed (0.5–2.0) + seed (0 = random) — see the CosyVoice Guide
    • Cloud config: API Key (Host-file-only, write-only input) + Project ID (optional) + test button + this month's local usage per tier (Standard 4M / Wavenet-Neural2-Chirp3 1M) — see the Cloud Guide
    • Voice packs: one-click install of voices from a registry
    • Preview: type text and press the play (triangle) button — a spinning loader shows while it is synthesizing/playing (click again to stop), failures show an inline message.
  4. RVC custom voices: read with your own trained RVC models, computed locally (index-free mode, advanced params — see the RVC guide).
  5. Gapless long reads: adaptive chunked progressive playback — probe-calibrated chunk size, play-while-converting, Web Audio sample-accurate joins, no gaps between chunks (see the design doc).
  6. Mini player while reading: pause / resume + playback speed (1x / 1.25x / 1.5x) on the message's action row; chunked long reads surface a visible "chunk x/y" counter.
  7. Themed tooltips & RVC onboarding: hover tooltips use the app's theme tokens (--dsw-*); the RVC panel opens with a first-time 3-step guide (per-OS startup commands + one-click diagnostics).
  8. Download audio: a download button on each message saves the synthesized audio — MP3 for Edge/Cloud reads, WAV for local (RVC/Index-TTS2/CosyVoice) reads — reusing the in-session cache so a just-read message downloads instantly. Long chunked reads have no single output file yet and report "not supported" instead of starting a wasteful re-synthesis.
  9. Read selected text: selecting text in a message shows a floating "朗读选中" chip — click it to read just that selection.
  10. Streaming long reads (Edge too): long plain-Edge reads also stream progressively (first chunk plays while the rest synthesize), reusing the gapless chunked pipeline — no more waiting for full synthesis.
  11. Approval voice alerts: optional voice broadcast of Agent approval events. When enabled, an approval request (approval/asked) is announced aloud and interrupts the current read (the agent is waiting on the decision); an approval result (approval/decided) is announced only when idle. Alerts always use Edge TTS with their own alert voice (independent of the RVC service), are deduplicated by approval id, and off by default. (Scope note: task-completion announcements are deferred to a later phase — the jobs subsystem has no session-event channel this plugin can observe yet.)
  12. Clean reads, quiet RVC: message Markdown (code blocks, tables, links, quotes, task lists) is cleaned before reading — no symbol soup; an opt-in "read raw Markdown symbols" toggle keeps verbatim reads for source code. With the Edge provider the RVC panel stays fully hidden behind a single "need a cloned voice?" entry, so casual reading is zero-config.

Requirements

  • DeepSeek Harness web profile (dsh web)
  • Node.js >= 22 (the worker uses the native WebSocket)
  • For local-voice providers only (nothing to install for Edge/Cloud):
    • RVC custom voices: a local RVC inference environment (an RVC WebUI or the portable runtime) and a running rvc-server.py — see the RVC Guide → "启动本地 RVC 服务" and the User Guide §4.2. macOS users start there too.
    • Index-TTS2 voices: the local index-tts2-nvidia bundle's API service (启动api服务.bat, default port 7880) — see the Index-TTS2 Guide.
    • CosyVoice clones: the local CosyVoice-2.0-0.5B bundle's cosy-server.py (启动cosy服务.bat, default port 7890) — see the CosyVoice Guide.

Install

Four equivalent ways (pick any one):

# dsh-market UI: search the market for the plugin name and install
# npm package:
dsh plugin --profile web add "@memef1f1y/dsh-plugin-tts"
# GitHub source:
dsh plugin --profile web add "github:1624318455/dsh-plugin-tts#main"
# or local development:
dsh plugin --profile web add "file:/path/to/dsh-plugin-tts/plugin"

Restart dsh web; the plugin then loads automatically as a profile bundle. Uninstalling any one of the above first is required before installing another: they all register the same loader entry id (tts), and the installer refuses duplicates (it rolls back instead of overwriting).

Voices (live-verified, Edge TTS)

Region Voices
Simplified Chinese Xiaoxuan 晓萱 · Xiaoyi 晓伊 · Yunxi 云希 · Yunyang 云扬 · Xiaoxiao 晓晓 · Yunjian 云健 · Yunxia 云夏 · liaoning-Xiaobei 晓北 · shaanxi-Xiaoni 晓妮
Taiwan HsiaoChen 曉臻 · HsiaoYu 曉雨 · YunJhe 雲哲
Hong Kong HiuGaai 曉佳 · HiuMaan 曉曼 · WanLung 雲龍
English Aria · Jenny · Guy · Sonia (UK)
Other Nanami 七海 (ja-JP) · SunHi (ko-KR) · Denise (fr-FR)

Note: legacy voices such as Xiaohan / Xiaomeng / Xiaorui / Xiaoshuang were removed by the Edge endpoint (1007 Unsupported voice) and are not listed.

Architecture

Layer Location Role
Host lib/index.mjs Registers webServer routes: /dsh-tts-api/speak (synthesis / chunk queue, all providers), /dsh-tts-audio/<id> (audio bytes), /dsh-tts-api/rvc-next (next chunk / job cancel), /dsh-tts-api/rvc-files + /rvc-compact-index + /rvc-packs* (RVC service proxy / packs), /dsh-tts-api/index-* and /dsh-tts-api/cosy-* (local service proxy + config), /dsh-tts-api/cloud-* (Cloud key/usage/test), /dsh-tts-api/notify (approval-alert queue), /dsh-tts-api/diagnose (one-click diagnostics); runs a zero-dependency Edge worker via node -e
Client lib/client.js Web Audio chunk player (sample-accurate joins, <audio> fallback) in shell.overlay + the UI entries (read-aloud button / auto-read toggle / settings panel); talks to the Host through fetch

The TTS worker mirrors node-edge-tts@1.2.10: Sec-MS-GEC query params (ticks rounded to the 5-minute boundary), Sec-MS-GEC-Version=1-143.0.3650.75, Path:audio binary framing, xml:lang derived from the voice locale, one retry on abnormal (1006) closures. Audio is audio-24khz-48kbitrate-mono-mp3.

Edge cases handled

  • Clicking the read button of the message being auto-read stops it; another message's button switches to manual reading.
  • Disabling auto-read never interrupts a manual read; it stops auto reads.
  • A newly completed message (auto on) interrupts the current read; text-less messages are skipped; session switches only stop auto reads.
  • Stopping / switching messages eagerly cancels the active RVC chunked job on the Host, so the local conversion service stops scheduling new chunks and releases GPU/memory promptly (no waiting for the lazy GC).
  • Repeatedly reading the same text + voice reuses the in-session audio cache (no re-synthesis); if the cached backing file was cleaned by the OS, it transparently re-synthesizes instead of serving a stale 404 URL.
  • If an Edge voice was removed by the endpoint (1007 Unsupported voice), the voice is pruned from the picker and the plugin auto-falls back to the default.
  • Audio is autoplay-unlocked on the first user gesture (Web Audio context resumed + silent clip) so reads aren't silently blocked by browser policy.
  • Esc / S (outside an input) stops the current read-aloud.
  • Synthesis / playback failures surface a themed toast (message read-aloud and auto-read no longer fail silently; the preview panel keeps its inline error). In RVC mode the error toast carries a one-click "read with Edge TTS instead" action; with the opt-in toggle "auto-switch to Edge TTS when RVC fails" in the RVC settings (off by default — RVC is fully local, auto-fallback sends the text to Microsoft's endpoint), a failed RVC read silently retries with the RVC base voice via Edge TTS and shows a warn toast.
  • Smart chunking never splits URLs / emails / decimals / versions ("3.14") mid-token — hard cuts slide to word/punctuation boundaries, and a tiny trailing sentence merges into the previous chunk instead of reading like a stutter.
  • Approval voice alerts (opt-in): the Host subscribes to the session/event firehose (approval/asked / approval/decided), dedupes by approval id, and the client polls /dsh-tts-api/notify?s=N — the first poll baseline-syncs the cursor so a page refresh never replays stale alerts; announcements are silent on failure (no error toasts while the agent loops).

Settings persistence

Layered storage:

  • Host file (~/.dsh/tts-rvc/settings.json, { version: 1, rvc: …, index: …, cosy: …, cloud: … }): service-class config — RVC service URL/model/index paths, Index-TTS2 and CosyVoice service URLs, Cloud API key (+ optional project ID). Shared across browsers, survives browser-data clears and incognito. Served/updated through GET /dsh-tts-api/rvc-config + POST /dsh-tts-api/rvc-config-save (and the matching index-config, cosy-config, cloud-config routes; legacy localStorage values migrate up once, then the file wins).
  • localStorage (dsh-tts-settings): UI prefs only — voice, auto-read toggle, provider, sound tuning, the raw-Markdown toggle, approval-alert and remaining RVC prefs. Never stores the service URL / model / index.

A "Reset to defaults" button in the settings panel restores defaults and clears both layers (best-effort on the Host file).

Custom voice (RVC)

Use your locally trained RVC model for voice conversion: switch the TTS provider to "自定义音色(RVC)" in the settings panel. First-time RVC users need two things: a model file (.pth) and a running local RVC service — see the RVC Guide or User Guide §4.2 for macOS/Windows/Linux startup commands. The full story — service startup, panel config, gapless chunked playback, compact index, voice-pack registry install, portable runtime, settings reference and troubleshooting — lives in the RVC Custom Voice Guide.

Public pack registry example: rvc-for-tts (设置 → 语音 → 音色包 → registry URL: https://raw.githubusercontent.com/1624318455/rvc-for-tts/main).

Troubleshooting (Edge TTS)

  • 403 / Sec-MS-GEC rejected: the Edge endpoint protocol or version check changed; update CHROMIUM_FULL_VERSION / TRUSTED_CLIENT_TOKEN inside the worker in lib/index.mjs.
  • 1007 Unsupported voice: the selected voice was removed from the endpoint; pick one from the table above.
  • No sound: check system volume, the browser autoplay policy (interact with the page once), or the synthesis logs ([tts] errors in the dsh web console).

RVC-specific troubleshooting: RVC Guide → Troubleshooting.

FAQ

Q: Is it big?

  • Default experience (Edge TTS): the plugin itself is tiny (MB scale) — no service or model to download.
  • You only need a local RVC portable package if you want custom voices (RVC). Its size mostly comes from the bundled offline Python runtime + inference deps + pretrained models:
    Platform Archive Extracted
    macOS (Apple Silicon) ~660 MB ~1.3 GB
    Windows (CPU-minimal) ~1–2 GB ~2–3 GB
    Windows (with NVIDIA GPU) ~6 GB ~7 GB
  • These are self-contained runtimes; the full RVC WebUI is 7.8 GB — we don't ship the WebUI / training / realtime parts this plugin doesn't need.
  • Index-TTS2 / CosyVoice also ship as local bundles (sizes vary by release — see the Index-TTS2 Guide / CosyVoice Guide preparation sections).

Q: Are the dependencies big?

  • The runtime is heavy but you don't install anything: the portable package bundles Python, ffmpeg (Windows) / PyAV (macOS), and all inference deps. Unzip and run — no compile, no env setup.
  • The plugin itself has minimal deps and only loads what it actually uses.

Q: Do I need a local TTS model?

  • With the default Edge TTS or Cloud TTS: no — they synthesize online.
  • With custom voices (RVC): you need your own RVC voice model (.pth), optionally a .index; the pretrained hubert / rmvpe are already bundled — you only provide your trained model.
  • With Index-TTS2 / CosyVoice: no training needed — you provide reference/prompt audio (a few seconds of clear speech) to the local bundle and pick it in the panel; see the Index-TTS2 Guide and CosyVoice Guide.

Q: Is it easy to install?

  • Install the plugin normally, then for RVC: download the portable package for your platform → unzip → run the launcher (double-click .command on macOS, .bat on Windows) → drop your .pth into assets/weights → pick it in the plugin panel.
  • Index-TTS2 / CosyVoice follow the same three steps (bundle → keep its service running → pick a voice in the panel); Cloud TTS only needs an API key pasted into settings. Each provider guide walks through it.
  • No compile, no manual Python/ffmpeg install. Note: the first macOS launch is slow (tens of seconds) while macOS scans the extracted runtimes (one-time); later launches take a few seconds.

Q: Do I need a paid API?

  • No for Edge TTS (free, no API key) and RVC / Index-TTS2 / CosyVoice (fully local and free — but each needs its local service package running, see Requirements above).
  • Google Cloud TTS is pay-as-you-go and needs your own API key: Chinese Standard voices are free for the first 4M chars/month, Wavenet/Neural2/Chirp for the first 1M chars/month (local counters; Google Cloud Billing is authoritative). See the Cloud Guide.
  • Note: Edge TTS is Microsoft's public client-side, free capability — fine for personal use; check Microsoft's terms for commercial / heavy usage.

Q: Did this modify the DSH core?

  • No. It's an independent plugin loaded through dsh's plugin mechanism. It doesn't touch the DSH main program and can be installed / paused / uninstalled anytime without affecting DSH or other plugins.

Other common questions

  • Do I need a GPU? No — CPU works. For speed you can use Apple Silicon MPS (macOS) or an NVIDIA GPU (Windows, which needs the bigger CUDA torch build).
  • Privacy? RVC conversion is fully local — audio never leaves your machine. Edge TTS sends the text-to-read to Microsoft's endpoint for online synthesis (assess which voice before choosing).
  • Apple Silicon only? The macOS build is arm64 (M1–M5); Intel Macs would need a separate x86_64 build.

UI language (i18n)

The settings panel has an Interface language selector at the top: Auto (follow browser) / 中文 / English.

  • Default "Auto" follows the browser/system language (Simplified Chinese and others → Chinese, everything else → English).
  • Switching applies immediately and is persisted to localStorage (dsh-tts-lang), surviving page reloads.
  • Covers the whole settings panel, bubble/read-aloud buttons, diagnostics, voice-pack panel, plus RVC service errors/progress hints.

Development

node tests/smoke.mjs   # fake-ctx route registration + real Edge TTS synthesis + audio serve assertions
npm run test:all       # full: smoke + live + patch + i18n + client-load

Hot-reload after editing lib/ (on Windows a file: install is a COPY, not a symlink, so the running dsh reads the profile copy):

Copy-Item lib/* $env:USERPROFILE\.dsh\profiles\web\node_modules\@memef1f1y\dsh-plugin-tts\lib\ -Recurse -Force
# then refresh the browser (bundles are re-read from disk per request; never use pnpm install --force)

Known limits

  • Voice / auto-read toggle / provider / RVC settings are persisted in layers (Host file for service config, localStorage for UI prefs — see "Settings persistence" above) and survive refresh (see "Settings persistence" above); the audio cache itself is in-session only (files live in the OS temp dir, cleaned by the OS), so a full restart re-synthesizes the first read of each text.
  • Synthesized audio is written to the OS temp dir and cleaned by the OS.
  • zh/en layout/visual fitting (English text is longer; may wrap/overflow; theme vars --dsw-*) must be eyeballed in the real dsh UI with the plugin loaded — this plugin ships no standalone HTML (its UI is slot-injected by the dsh web host), so it cannot be headless-screenshotted here (tests/client-load.mjs asserts the in-memory render only, not real DOM/CSS).

License

MIT

Content from the project README on GitHub ↗

Links

More in this category

View the whole category →

Community comments

Comments are public GitHub Discussions. Loading them connects to GitHub and Giscus; a GitHub account is required to post.