DeepSeek Harness Plugin

WayneYu430/dsh-voice-agent#voice-app

Stars ★ 2 Downloads (30d) 679 Category Voice & Audio Added 2026-08-27 npm @wayneyu430227/dsh-voice-agent

A conversational voice frontend Agent for dsh: speak naturally over ByteDance Duplex, delegate requests to background tasks, and hear their asynchronous results reported by voice.

Install

# from npm (prebuilt)

dsh plugin --profile web add @wayneyu430227/dsh-voice-agent

# from GitHub (first run asks for allowBuilds approval — follow the hint, retry)

dsh plugin --profile web add github:WayneYu430/dsh-voice-agent#path:/packages/voice-app

Any plugin you install runs third-party code with your own permissions — it can read your files, use your credentials, and reach the network, and tool approvals don’t sandbox it. GitHub-sourced plugins also run build scripts at install time — pnpm blocks those until you allow them, so an install can stop with ERR_PNPM_GIT_DEP_PREPARE_NOT_ALLOWED or ERR_PNPM_IGNORED_BUILDS; dsh prints the exact key to add under allowBuilds in your profile’s pnpm-workspace.yaml, and the install works on the next run. Allowing a build is a trust decision: only install sources you trust, and pin a commit (github:owner/repo#sha).

README

English | 中文

Patch-layer bundle for the voice profile. It layers over dsh-base and dsh-web-app, then mounts the provider-neutral voice seam, ByteDance Duplex provider, ordinary text-Agent consumer, dedicated browser WebSocket, and microphone/playback client surface. Starting a Voice conversation creates a fresh ordinary, ungrouped source Session at the current Workspace or Session directory and attaches the provider transport to it; keeping that source outside Workspace membership prevents standard blank-Session reuse. Each accepted delegation creates an independent ordinary task Session, linked from a compact card while root-owned audio continues across navigation. A plugin-owned browser history index exposes saved Voice Sessions from the sidebar without Host or Workspace-specific metadata. Configure the Duplex access key through the DUPLEX_API_KEY credential reference.

Model Experience

Voice profile composition

What the model sees

Voice-initiated work reaches a fresh task Agent only as an accepted realtime_delegation envelope and exact-id updates. That task Agent alone receives the scoped send_voice_message backend tool for STATUS and COMPLETE; the bridge creates the target directly, so no project-listing tool is added. The Duplex frontend Agent owns the spoken conversation and sees exactly its three orchestration tools.

Token effect

Accepted delegation text, ordinary task work, and backend reporting calls consume text-model tokens; Duplex separately spends provider tokens on the voice conversation and task summaries.

KV Cache effect

Only accepted commands extend the independent task Agent's history; the Voice Session transcript and frontend conversation do not alter its request prefix.

Known Limitations and Deferred Work

  • The shipped first provider is Duplex; the service seam is provider-neutral so Realtime/Live providers can be added without changing the assistant consumer.
  • Raw audio and provider conversation state remain process-local; completed or interrupted utterance text and task links are durable.
  • The filtered Voice history index is browser-local; clearing site data does not delete the underlying Sessions.
  • The browser client surface targets the dsh Web UI: it is emitted by the copied dsh client tsdown preset and loads through the dsh web runtime's window.__ModuleLoader__ contract. The server-side packages are transport-agnostic, but the microphone/playback UI is not a standalone browser plugin.

Content from the project README on GitHub ↗

Links

More in this category

View the whole category →

Community comments

Comments are public GitHub Discussions. Loading them connects to GitHub and Giscus; a GitHub account is required to post.