A conversational voice frontend Agent for dsh: speak naturally over ByteDance Duplex, delegate requests to background tasks, and hear their asynchronous results reported by voice.
Install
# from npm (prebuilt)
dsh plugin --profile web add @wayneyu430227/dsh-voice-agent
# from GitHub (first run asks for allowBuilds approval — follow the hint, retry)
dsh plugin --profile web add github:WayneYu430/dsh-voice-agent#path:/packages/voice-app
Any plugin you install runs third-party code with your own permissions — it can read your files, use your credentials, and reach the network, and tool approvals don’t sandbox it. GitHub-sourced plugins also run build scripts at install time — pnpm blocks those until you allow them, so an install can stop with ERR_PNPM_GIT_DEP_PREPARE_NOT_ALLOWED or ERR_PNPM_IGNORED_BUILDS; dsh prints the exact key to add under allowBuilds in your profile’s pnpm-workspace.yaml, and the install works on the next run. Allowing a build is a trust decision: only install sources you trust, and pin a commit (github:owner/repo#sha).
README
English | 中文
Patch-layer bundle for the voice profile. It layers over dsh-base and dsh-web-app, then mounts the provider-neutral voice seam, ByteDance Duplex provider, ordinary text-Agent consumer, dedicated browser WebSocket, and microphone/playback client surface. Starting a Voice conversation creates a fresh ordinary, ungrouped source Session at the current Workspace or Session directory and attaches the provider transport to it; keeping that source outside Workspace membership prevents standard blank-Session reuse. Each accepted delegation creates an independent ordinary task Session, linked from a compact card while root-owned audio continues across navigation. A plugin-owned browser history index exposes saved Voice Sessions from the sidebar without Host or Workspace-specific metadata. Configure the Duplex access key through the DUPLEX_API_KEY credential reference.
Model Experience
Voice profile composition
What the model sees
Voice-initiated work reaches a fresh task Agent only as an accepted realtime_delegation envelope and exact-id updates. That task Agent alone receives the scoped send_voice_message backend tool for STATUS and COMPLETE; the bridge creates the target directly, so no project-listing tool is added. The Duplex frontend Agent owns the spoken conversation and sees exactly its three orchestration tools.
Token effect
Accepted delegation text, ordinary task work, and backend reporting calls consume text-model tokens; Duplex separately spends provider tokens on the voice conversation and task summaries.
KV Cache effect
Only accepted commands extend the independent task Agent's history; the Voice Session transcript and frontend conversation do not alter its request prefix.
Known Limitations and Deferred Work
- The shipped first provider is Duplex; the service seam is provider-neutral so Realtime/Live providers can be added without changing the assistant consumer.
- Raw audio and provider conversation state remain process-local; completed or interrupted utterance text and task links are durable.
- The filtered Voice history index is browser-local; clearing site data does not delete the underlying Sessions.
- The browser client surface targets the dsh Web UI: it is emitted by the copied dsh client tsdown preset and loads through the dsh web runtime's
window.__ModuleLoader__contract. The server-side packages are transport-agnostic, but the microphone/playback UI is not a standalone browser plugin.
Links
More in this category
PolinniZhong/dsh-omi-voice★ 75
In-chat read-aloud for DeepSeek Harness: tap to read, pause and resume AI replies with natural Doubao TTS voices (BYOK), reading only the final answer with code, tables and diagrams filtered; local engine, plugin keeps no API key.
PensiveFei/dsh-voice-scribe★ 35
Voice input for the web UI: tap Alt (or Alt+Space) to start/stop dictation, browser Web Speech (zero-config) or OpenAI-compatible cloud ASR, optional polish through DSH-configured LLM, settings UI.
1624318455/dsh-plugin-tts★ 24
Reads assistant replies aloud via free Edge TTS or your own RVC voice models: read-aloud buttons + auto-read, adaptive chunked progressive playback (gapless long reads), one-click voice-pack installs from a registry, and a portable RVC runtime.
WizisCool/dsh-ears★ 22
Voice input plugin for DeepSeek Harness (dsh): a microphone button in the composer turns speech into a draft transcript, with a choice of speech-recognition backends, optional polish through dsh own LLM routes, and a native settings page.
ppy-web/dsh-plugin-xiaomi-mimo-tts★ 19
Adds Xiaomi MiMo text-to-speech to DSH Web with assistant-message read-aloud, PCM streaming, preset and custom voice design, browser speech fallback, playback controls, and optional UI sounds.
PerryLink/dsh-talk★ 17
Voice I/O for DeepSeek Harness — speech-to-text and text-to-speech over the microphone and audio output.
Community comments
Comments are public GitHub Discussions. Loading them connects to GitHub and Giscus; a GitHub account is required to post.