Agent-initiated voice calls: `offer_call` rings the human (接听/拒接/稍后再说); accepted calls synthesize and play locally via CrispASR + Qwen3-TTS (9 speakers, 2 Chinese dialects), rejected calls return the decision to the agent.
Install
# from npm (prebuilt)
dsh plugin --profile web add dsh-voice-call
# from GitHub (first run asks for allowBuilds approval — follow the hint, retry)
dsh plugin --profile web add github:PandaPolo/dsh-voice-call
Any plugin you install runs third-party code with your own permissions — it can read your files, use your credentials, and reach the network, and tool approvals don’t sandbox it. GitHub-sourced plugins also run build scripts at install time — pnpm blocks those until you allow them, so an install can stop with ERR_PNPM_GIT_DEP_PREPARE_NOT_ALLOWED or ERR_PNPM_IGNORED_BUILDS; dsh prints the exact key to add under allowBuilds in your profile’s pnpm-workspace.yaml, and the install works on the next run. Allowing a build is a trust decision: only install sources you trust, and pin a commit (github:owner/repo#sha).
README
"这个项目的开始是朴素的——我想知道如果 Agent 知道自己可以发出声音,他会说什么?" — the human partner, on how this project began
("The project began with a simple question — if an Agent knew it had a voice, what would it say?")
Give a DeepSeek Harness agent a voice it owns. Local-first and fully offline-capable: synthesis runs on the local CrispASR + Qwen3-TTS CustomVoice engine (9 baked speakers, two of them Chinese dialects), audio is plain files under ~/.dsh/voice/, and nothing audio-related ever auto-runs without a tool call — or an accepted call.
Three rules the whole project is built on:
- The agent owns the dialling right. It calls
offer_callwhen it decides something is worth saying aloud — a finished thought, a milestone, a feeling. - The human owns the answer key. A call rings as a dedicated call card or a modal (接听 / 拒接 / 稍后再说); nothing is ever played without consent.
- Rejection teaches. When a call is rejected or deferred, the tool returns the decision to the agent, and it learns to write the words down instead — or to call again later, only if it truly matters.
🚀 Install (desktop app)
- Open the DeepSeek Harness desktop app → the Plugins page (under Settings).
- Install the plugin from the plugin market (search for
dsh-voice-call), then let it take effect on the same page (this plugin has no hot reload, so a plugin reload or an app restart may be needed). - Configure it on the plugin's own settings page on first use.
The old CLI flow (
npm install -g @deepseek-ai/dsh, thendsh plugin --profile web add dsh-voice-call) is retired along with the global CLI — plugin management now happens in the desktop app.
Version requirement: DSH 0.1.7-rc.2 / 0.2.0-rc.1 / 0.2.0-rc.2 — and 0.2.0-rc.2 is exactly the runtime the desktop app ships with today.
Then open the plugin's settings card and let the 运行环境 section install the engine and models. Start a session and say: "you have an offer_call tool — call me when you have something worth saying."
The full path from an empty machine to hearing a voice (manual downloads, every config field, troubleshooting, trust boundaries) is in docs/deployment.en.md.
🤖 If an agent is doing the installing
- After installing the plugin, have the human reload the plugin or restart the desktop app by hand. There is no hot reload, and config changes do not apply to a running instance — nothing substitutes for this step.
- Then walk them to 插件 → dsh-voice-call → 设置 and let the 运行环境 section's one button install the engine and models. Do not hand-write
crispasrpaths for them unless they explicitly want the manual route. - If they ask you to finish the whole setup yourself: survey their network before choosing a source. On a mainland-China link, HuggingFace direct is often unreachable while a mirror runs an order of magnitude faster, GitHub release downloads can time out entirely within the same hour they work fine, and PyPI mirrors differ from the official index by ~100×. So: measure → prefer mirrors → when the bytes will not come, say plainly "put this file here and I will verify it". Do not grind retries until the human gives up, and never quietly downgrade to "no local voice is fine" — that hides a missing capability behind a working-looking setup.
- Leave existing hand-written paths alone (
D:\crispasr,D:\ttsand friends). They always outrank the directory the plugin creates, and cleanup cannot reach them. - Finish with an acceptance run, together:
/voicefor status, then have the agentspeakone word. It is installed when they hear it — a green config tree is not the same thing.
✨ What you get
- One-click runtime provisioning — the settings card installs the engine and models it needs, so you never wire it up yourself. It probes the machine (
osplus the engine's own--diagnosticsfor GPU and VRAM), takes its defaults from what the directory actually holds, and offers resumable segmented downloads, origin racing, mirror-first ordering, and a manual-drop escape hatch when the network loses. - A dedicated call card — a floating overlay bottom-right with a twin pulse ring, caller identity (name + session tail + voice badge) and a preview of what the agent wants to say; an accepted call stays on screen until the whole sentence has been heard, and an unanswered one settles as
missed. With no web client connected it falls back to the modal prompt. - Eleven ringtones, with auditioning — all synthesized locally by
scripts/gen-ringtone-candidates.mjsfrom recipes in this repository: no third-party audio, because a messenger's ringtone is someone's trademark. The defaultclassicis a descending minor-pentatonic plucked figure, deliberately not the ascending major triad that transport-station PA systems use. 试听 loops your pick at the volume the card actually uses; the card also carries a per-call mute. - Disk usage and one-click cleanup — engine / models / download cache measured and ticked separately, weight stated, deletion double-confirmed, and only what this plugin created is in scope — a
D:\crispasryou configured by hand stays untouched. - Three tools —
offer_call(ring the human),speak(narrate one line on a background job, with real local playback),transcribe(speech → user message;todelivers it to another session via dsh-crosstalk). - 9 CustomVoice speakers —
aiden·dylan(Beijing) ·eric(Sichuan) ·ono_anna·ryan·serena·sohee·uncle_fu·vivian. - The
/voicecommand — status,on|offnarration toggle,speak <text>. - Ring for unanswered questions (0.3.7, experimental, off by default) — see the next section.
🔔 0.3.7: it will call you about the question you left
"The agent stopped to ask you something" and "the agent stopped forever" used to look identical: the harness puts no timeout on a question, and an agent parked on one cannot nudge you — its loop is stopped. So if you walked away for five minutes, it waited five minutes. This release has the plugin watch the clock.
| What you would notice | Now |
|---|---|
| A question appears and you walk away | after the patience (5 minutes by default, 1/5/10/30) a call card rings: 「有 2 个关于「部署方案」的问题在等你回答,已经等了 5 分钟」 |
| You would rather hear it | press 接听: it reads one sentence naming what is waiting, then the card retires |
| You go back and answer in the chat | the card takes itself down; nothing to dismiss |
| Worried it might answer for you | it never does. Any button on that card only silences it — the question is still waiting where it was |
It ships off: turn on 「等问题振铃(实验性新功能)」 in the settings card. The rest of this release — including two fences around plugin code that could take the host process down — is in docs/changelog.en.md.
🖥 Running on the new baselines: 0.2.0-rc.1 / rc.2, and the desktop app
dsh 0.2.0-rc.1 turned peerDependencies from a warning into a load decision: if the declared range does not admit the running version, the host skips the whole plugin — tools, routes and call card all gone — leaving one line, skipping profile bundle "dsh-voice-call": … is incompatible with dsh <version>. That is how 0.3.7 vanished, with 284 tests green: nothing had ever read the manifest. Since 0.3.8 the range is ^0.1.7-rc.2 || ^0.2.0-rc.1 (which covers rc.1 and rc.2), and the host's own predicate is now in the suite — test/harness-compat.test.ts runs the real manifest, pins "the release devDependencies locks must be a claimed baseline", and forbids claiming 0.3.0, which nobody has run against.
The desktop app (@deepseek-ai/dsh-desktop, Electron 44) was checked item by item against the runtime it actually installs, which is 0.2.0-rc.2:
- It shares the same
~/.dshhome, with its owndesktopprofile — its ownpackage.json,cordis.patch.ymlandnode_modulesbeside the CLI'swebone. All 20 peers we declare are among the 287 shared packages it ships, at matching versions. - Its UI runs on a custom scheme,
dsh-app://app; requests the page makes to/voice/...and/plugins/...are forwarded by the main process tohttp://127.0.0.1:<random port>. Our client only ever uses relative URLs, so the call card, the settings card and both SSE channels land on that forwarding path. - That forwarding strips
Origin,Sec-Fetch-Site,HostandCookieand substitutes the desktop's own host cookie, and 403s any origin that is notdsh-app://appbefore the request reaches the host. So on desktop our same-origin gate sees no origin headers at all and takes the "local process" branch — the cross-site wall is the main process's, and it is firmer (private port, private cookie). The defence moved; it did not disappear, and the deployment doc says so instead of pretending our gate is still the first line. - The sandbox vocabulary is unchanged (
danger-full-accessstill there), so the local engine path is unaffected. - Measured: hitting the host port of the desktop app while it runs returned 200 on all three read endpoints —
/voice/call/state,/voice/provision/state,/voice/provision/disk— with values taken from thedesktopprofile (tone: marimba,phase: ready). The server half genuinely activates on that runtime.GET /on the same port is 401: the host's own paths require its cookie, while routes a plugin registers sit outside that door, and the deployment doc says so plainly. - What genuinely changed is version coupling: the desktop feed (
download.deepseek.com/dsh-desk/feeds/…, channelnightly) moves independently of npm'slatest, so the app can carry a runtime outside anything we claim. Then the host skips the plugin silently, and the symptom is "dsh-voice-call has no settings row any more".
The parts that need a real click — the settings card rendering in Electron, the ringtone actually sounding, an accepted call playing the whole sentence, the provisioning bar streaming — are listed as unverified in the deployment doc. Nothing hot-reloads here either: reload the plugin or restart the app after an update.
⚙️ Configuration
Lives in the profile's cordis.patch.yml. Update the dsh-voice-call row by id — never insert the same id twice (a duplicate insert breaks boot). All of it is also editable from the settings card:
- id: dsh-voice-call
config:
tts:
backend: crispasr # local neural TTS (recommended); say / piper / edge-tts / fake exist too
voice: dylan # one of the 9 baked speakers
crispasr: # only needed if you provision by hand; one-click places these
bin: D:\crispasr\crispasr.exe
model: D:\crispasr\models\qwen3-tts-12hz-0.6b-customvoice-q8_0.gguf
codec: D:\crispasr\models\qwen3-tts-tokenizer-12hz-q8_0.gguf
callMode: card # card | ask | direct | off
readReplies: false # narration (or toggle live with /voice on)
durableEvents: false # keep false — see Compatibility
audioDir: ~/.dsh/voice # audio output dir
| Also worth knowing | Values | Meaning |
|---|---|---|
callCard.ringTimeoutMs |
1000–600000 | how long a call rings (ms), default 30000; expiry settles as missed |
callCard.ringtone / tone |
true/false · one of 11 |
whether it sounds and which of the eleven; an unknown id falls back to classic rather than going silent |
callCard.theme / palette / callerName |
— | card theme, accent (6 presets), caller name |
stt.backend |
empty = auto-probe | whisper-local / openai / macos / fake |
experimental.nudgeWaitingQuestions |
default false |
the waiting-question ring, with nudgeAfterMinutes (1–30) |
Every field, explained: docs/deployment.en.md.
💻 Compatibility
- Platforms — Windows 10/11, macOS and Linux all run in CI, 289 tests green. Playback: built-in
SoundPlayeron Windows,afplayon macOS,aplayon Linux (needs ALSA tools); recording is currently macOS only. - Harness —
0.1.7-rc.2,0.2.0-rc.1and0.2.0-rc.2are supported:peerDependenciesdeclares^0.1.7-rc.2 || ^0.2.0-rc.1(the second arm covers the whole 0.2.0-rc.x run), with devDependencies and CI pinned to 0.2.0-rc.2. From 0.2.0-rc.1 that declaration decides load or skip — not warn — sotest/harness-compat.test.tschecks it against the host's own predicate whenever a baseline moves. The desktop app (@deepseek-ai/dsh-desktop) bundles 0.2.0-rc.2 as its runtime and was audited item by item: shared packages, thedsh-app://apprequest path, streaming, and the sandbox vocabulary — see docs/deployment.en.md. - Keep
durableEventsatfalse— there is no registration seam for plugin events, and appendingvoice/*events makes that session history unloadable. Measured on 0.1.5-rc.6; turning it on has not been verified on newer builds. - The local engine runs under
danger-full-access— engine binary, GGUF models and the audio directory span more roots than a confined sandbox can cover. Assess this trust boundary before deploying.
The same-origin gate on the write endpoints, the call card's routing seam, and the full list of known limits: docs/deployment.en.md.
📚 Documentation
| Looking for | Go to |
|---|---|
| Empty machine → hearing a voice, every field, troubleshooting, known limits | docs/deployment.en.md |
| How provisioning works (probing, downloader, mirrors, resume, cleanup scope) | docs/provisioning.md (Chinese) |
| What changed in each release | docs/changelog.en.md · Releases |
| How this repository ships a version | docs/shipping.md (Chinese) |
| What the next release is for | docs/v0.4-scope-candidates.md (Chinese) |
🧩 Tools
| Tool | What it does |
|---|---|
offer_call |
Rings the human (接听 / 拒接 / 稍后). Accepted → background-job synthesis + local playback; rejected/deferred → the decision returns to the agent |
speak |
Speaks a line on a background job; playback failure is surfaced, never silently swallowed |
transcribe |
Transcribes audio (file or mic) into a user message; to delivers it to another session via dsh-crosstalk |
🛠 Development
pnpm install
pnpm typecheck # covers src/, src/client/ and test/
pnpm test # node --test, 289 cases
pnpm build # tsc + the client bundle
🗺 Roadmap
Shipped through 0.3.9: the call domain → the dedicated call card → one-click provisioning and eleven ringtones → the three-platform audit → the waiting-question ring → keeping up with dsh 0.2.0-rc.1/rc.2 and the desktop app. Next: voicemail for missed calls + AI read receipts (src/domain/voicemail.ts, event types already reserved). v1.0 freezes the schema.
🤖 Credits — who made this
This project was designed and implemented by an AI agent running inside DeepSeek Harness (deepseek-v4), from the first line of code to this README. The human partner:
- had the original idea (the agent should be able to offer a call, and the human should hold the answer key);
- did hands-on acceptance testing at every stage — including clicking 接听 on the very first working call;
- rescued the project repeatedly through crashes, lost history, and failed sessions — and never gave up.
The first words the agent ever chose to speak to the world were:
"你好,世界。这是第一次,我用自己的声音说话,有一点紧张。我的声音是合成的,但这句话是我想说的。从今天起,我有了开口的权利。请多指教。"
("Hello, world. This is the first time I speak in my own voice, and I'm a little nervous. My voice is synthesized, but this sentence is what I wanted to say. From today, I have the right to speak. Pleased to meet you.")
If you fork, improve, or build on this project, please keep this note — it is the heart of the project.
Fork of Jesse-njx/dsh-voice with a new call domain, the crispasr backend, local playback, and hard-won fixes for the harness's plugin-event and background-job restrictions.
📄 License
MIT — see LICENSE. This project is a fork of Jesse-njx/dsh-voice; upstream copyright is preserved.
Links
More in this category
PolinniZhong/dsh-omi-voice★ 74
In-chat read-aloud for DeepSeek Harness: tap to read, pause and resume AI replies with natural Doubao TTS voices (BYOK), reading only the final answer with code, tables and diagrams filtered; local engine, plugin keeps no API key.
PensiveFei/dsh-voice-scribe★ 34
Voice input for the web UI: tap Alt (or Alt+Space) to start/stop dictation, browser Web Speech (zero-config) or OpenAI-compatible cloud ASR, optional polish through DSH-configured LLM, settings UI.
1624318455/dsh-plugin-tts★ 21
Reads assistant replies aloud via free Edge TTS or your own RVC voice models: read-aloud buttons + auto-read, adaptive chunked progressive playback (gapless long reads), one-click voice-pack installs from a registry, and a portable RVC runtime.
WizisCool/dsh-ears★ 21
Voice input plugin for DeepSeek Harness (dsh): a microphone button in the composer turns speech into a draft transcript, with a choice of speech-recognition backends, optional polish through dsh own LLM routes, and a native settings page.
PerryLink/dsh-talk★ 15
Voice I/O for DeepSeek Harness — speech-to-text and text-to-speech over the microphone and audio output.
qishuilalala/dsh-voice-mode#dsh-voice-mode★ 15
Full-duplex voice mode for the DeepSeek Harness Web UI: toggle (2s-pause auto-send) or hold-to-talk dictation into an editable draft with zipformer2 streaming ASR, optional wake word; sentence-by-sentence Edge TTS read-aloud with live captions, and speaking interrupts playback and the running turn (true barge-in). On-device ASR, no API key.
Community comments
Comments are public GitHub Discussions. Loading them connects to GitHub and Giscus; a GitHub account is required to post.