Reads assistant replies aloud via Fish Audio API only (bring your own key): per-message read-aloud, auto-read toggle, and a settings page for model, voice reference_id, encrypted API key, and proxy.
Install
# from npm (prebuilt)
dsh plugin --profile web add dsh-fish-tts
# from GitHub (first run asks for allowBuilds approval — follow the hint, retry)
dsh plugin --profile web add github:MaRi23333/dsh-fish-tts
Any plugin you install runs third-party code with your own permissions — it can read your files, use your credentials, and reach the network, and tool approvals don’t sandbox it. GitHub-sourced plugins also run build scripts at install time — pnpm blocks those until you allow them, so an install can stop with ERR_PNPM_GIT_DEP_PREPARE_NOT_ALLOWED or ERR_PNPM_IGNORED_BUILDS; dsh prints the exact key to add under allowBuilds in your profile’s pnpm-workspace.yaml, and the install works on the next run. Allowing a build is a trust decision: only install sources you trust, and pin a commit (github:owner/repo#sha).
README
English | 中文
A third-party text-to-speech (TTS) plugin for the
DeepSeek Harness (DSH) Web GUI:
one-click read-aloud for every assistant reply, an auto-read toggle in the
composer, and configurable model / voice / encrypted API key / proxy. Fish Audio
API only — bring your own API key. Works with any voice id (reference_id) you
are authorized to use, including voices you cloned on Fish Audio. The UI is
bilingual (English / 中文, follows the DSH locale).
30-second comparison vs. typical Edge TTS plugins
| dsh-fish-tts (this plugin) | typical Edge TTS plugins | |
|---|---|---|
| Engine | Fish Audio official API (only; bring your own key) | Microsoft Edge built-in voices |
| Voice | Your own reference_id (incl. voices you cloned, must be authorized) |
Fixed Edge voice library |
| API key | Required (AES-256-GCM encrypted in the settings page) | None |
| Best for | Users with a Fish Audio account who want their own or cloned voices | Quick free trials with fixed voices |
Features
- Read-aloud action: every finalized assistant message gets a speaker button in its action strip (same icon style as the native actions). Click to synthesize and play that reply; click the same message's button again while playing to stop (no restart from the top), clicking another message's button switches playback straight over, and a second click while synthesis is still in flight cancels it. Markdown is cleaned before speaking: paths, URLs, long ids and code blocks are replaced with placeholders instead of being read out.
- Auto-read: a small speaker toggle in the composer tool row (synced with the settings page). When enabled, replies that arrive after the page loaded are read automatically.
- Settings page (Settings → Voice (Fish TTS)):
- TTS model (datalist suggestions + free text; e.g. s2.1-pro-free / s2.1-pro / s2-pro; saved values apply immediately; default s2.1-pro-free)
- Voice
reference_id(required — voices are personal data, the plugin ships no default; synthesis is refused with a hint while empty) - API key (AES-256-GCM encrypted in
$DSH_HOME/fish-tts/settings.jsonon this machine;key.binis generated once and ACL-tightened on Windows; the key never appears in any GET response, log line or the repository) - HTTP proxy (e.g.
http://127.0.0.1:7890, leave empty for direct) - Test clip, auto-read toggle, volume slider (default 60%), playback-speed slider (0.5–2.0×, pitch-preserving; fixed at 1× where the browser lacks support)
Screenshots
Install
One command, from npm (recommended):
npx @deepseek-ai/dsh plugin --profile web add dsh-fish-tts
Then restart dsh web (stop the process, run dsh web again), refresh the page, and open Settings → Voice (Fish TTS).
Other install sources:
# From GitHub (git-hosted plugins build on install)
npx @deepseek-ai/dsh plugin --profile web add github:MaRi23333/dsh-fish-tts
# From a local checkout
git clone https://github.com/MaRi23333/dsh-fish-tts.git
cd dsh-fish-tts
pnpm install && pnpm run build
npx @deepseek-ai/dsh plugin --profile web add /absolute/path/to/dsh-fish-tts
The repo commits
lib/build artifacts, so git installs need no local build; after changing sources runpnpm run buildand restart.
Switching from the GitHub install to npm: a bare
add dsh-fish-ttsis a silent no-op when the git version is already installed (pnpm considers the same-name dependency satisfied). Usenpx @deepseek-ai/dsh plugin --profile web add dsh-fish-tts@latestinstead, orremovefirst and thenadd.
When installing a freshly published npm version, pnpm's supply-chain protection may automatically add
minimumReleaseAgeExclude: [dsh-fish-tts@…]to the profile'spnpm-workspace.yaml. This is expected and harmless.
Verify the install
- Open Settings → Voice (Fish TTS);
- Fill in your API key and voice
reference_id; - Click Save settings (the API key status turns "configured");
- Click Test — hearing the test sentence in your voice means the install works.
The Test button performs one real synthesis and verifies the key, the voice and the proxy configuration in a single click. It stays disabled while the settings are unsaved or the voice is empty.
Configuration
First run: open Settings → Voice (Fish TTS), fill in model, voice, API key (from Fish Audio) and a proxy if needed, save, then use the Test button. All settings take effect immediately after saving — no restart required.
You may also add a config to the fish-tts row in your profile's cordis.patch.yml
(settings-page values take precedence):
- id: fish-tts
config:
model: s2.1-pro-free
format: wav
stateDir: /custom/state/dir
Config keys
| Key | Default | Description |
|---|---|---|
model |
'' |
Default model (settings-page value wins) |
voice |
'' |
Default voice reference_id (settings-page value wins) |
format |
wav |
wav / mp3 / opus / pcm |
apiKey |
'' |
Usually empty; the encrypted settings-page key wins, then env FISH_API_KEY |
apiKeyFile |
'' |
Read FISH_API_KEY from a dotenv file |
proxy |
'' |
HTTP(S) proxy (settings-page value wins) |
stateDir |
$DSH_HOME/fish-tts |
Settings / key-file directory |
Security
- The API key is persisted only in encrypted form (AES-256-GCM, per-machine random
key.bin, 0600/ACL tightened) and never written to the repo, logs or any GET response. - Write routes (synthesize/config) require
application/jsonand validate same-origin/loopbackOrigin, blocking cross-site form abuse. - Local-only: every
/fish-tts/*route rejects non-loopback peers (127.0.0.1 / ::1 / ::ffff:127.0.0.1) with 403, even if the host listens on 0.0.0.0. - Proxy URLs with username/password are refused on save; credentialed
HTTPS_PROXY/HTTP_PROXYenv vars are likewise ignored (no leak, no fallback) — use a credential-less proxy or direct connection. - Proxy addresses, models and voices are machine-local user settings; the repo carries no personal data.
- Synthesis text is capped at 12000 characters; results are cached in-process (max 200 entries), cleared on restart.
Develop
pnpm install
pnpm run typecheck
pnpm run test # node:test suite (upstream Fish API is locally mocked, no network)
pnpm run build # host: lib/index.js; client: lib/client.js (ModuleLoader CJS closure)
pnpm run smoke # host entry + client ModuleLoader smoke tests
pnpm run check:pack # npm pack content whitelist check
Requires Node >= 22 (Node 20 is EOL). CI (
.github/workflows/ci.yml) runs the full gate chain on Node 22 and 24 and verifieslib/artifacts match the committed ones.
- Host side lives in
src/index.ts(Node; registers the/fish-tts/*routes and the settings store). - Client side lives in
src/client/(React; registers theconversation.chat.assistant-actions,conversation.input.leftandsettings.sectionslots). - Adapted and verified for the DSH
0.1.2-rc.1session API and UI icon changes; if the API drifts on other versions, align with the matching tag of the deepseek-harness repo.
License
Compliance & Third-Party Notice
- This is a third-party open-source plugin, not affiliated with, sponsored, or endorsed by Fish Audio / Hanabi AI Inc. "Fish Audio" is a trademark of its owner and is used here descriptively only.
- The plugin does not distribute or host API keys — use your own Fish Audio account and key, and keep it safe.
- Fish Audio's free tier is for personal, non-commercial use only; commercial use requires a paid plan. See the Terms of Use.
- Only use voices (reference_id) you are authorized to use. Do not clone or imitate the voice of public figures, celebrities, or private individuals without permission. See the Acceptable Use Policy.
- When distributing generated audio, disclose that it is AI-synthesized; do not mislead listeners into believing it is a real human recording.
- Using this plugin means you agree to Fish Audio's terms; the official pages prevail if updated.
Independent community project — not affiliated with or endorsed by DeepSeek.
Links
More in this category
PolinniZhong/dsh-omi-voice★ 74
In-chat read-aloud for DeepSeek Harness: tap to read, pause and resume AI replies with natural Doubao TTS voices (BYOK), reading only the final answer with code, tables and diagrams filtered; local engine, plugin keeps no API key.
PensiveFei/dsh-voice-scribe★ 34
Voice input for the web UI: tap Alt (or Alt+Space) to start/stop dictation, browser Web Speech (zero-config) or OpenAI-compatible cloud ASR, optional polish through DSH-configured LLM, settings UI.
1624318455/dsh-plugin-tts★ 21
Reads assistant replies aloud via free Edge TTS or your own RVC voice models: read-aloud buttons + auto-read, adaptive chunked progressive playback (gapless long reads), one-click voice-pack installs from a registry, and a portable RVC runtime.
WizisCool/dsh-ears★ 21
Voice input plugin for DeepSeek Harness (dsh): a microphone button in the composer turns speech into a draft transcript, with a choice of speech-recognition backends, optional polish through dsh own LLM routes, and a native settings page.
PerryLink/dsh-talk★ 15
Voice I/O for DeepSeek Harness — speech-to-text and text-to-speech over the microphone and audio output.
qishuilalala/dsh-voice-mode#dsh-voice-mode★ 15
Full-duplex voice mode for the DeepSeek Harness Web UI: toggle (2s-pause auto-send) or hold-to-talk dictation into an editable draft with zipformer2 streaming ASR, optional wake word; sentence-by-sentence Edge TTS read-aloud with live captions, and speaking interrupts playback and the running turn (true barge-in). On-device ASR, no API key.
Community comments
Comments are public GitHub Discussions. Loading them connects to GitHub and Giscus; a GitHub account is required to post.