DeepSeek Harness Plugin

DIAG5/dsh-better-input

Stars ★ 30 Downloads (30d) 2,194 Category UI Enhancements Added 2026-08-23 npm dsh-better-input

Input experience enhancement: voice input, AI polish, anti-overwrite, prompt optimization with diff preview, and auto locale switching following DSH UI language.

Install

# from npm (prebuilt)

dsh plugin --profile web add dsh-better-input

# from GitHub (first run asks for allowBuilds approval — follow the hint, retry)

dsh plugin --profile web add github:DIAG5/dsh-better-input

Any plugin you install runs third-party code with your own permissions — it can read your files, use your credentials, and reach the network, and tool approvals don’t sandbox it. GitHub-sourced plugins also run build scripts at install time — pnpm blocks those until you allow them, so an install can stop with ERR_PNPM_GIT_DEP_PREPARE_NOT_ALLOWED or ERR_PNPM_IGNORED_BUILDS; dsh prints the exact key to add under allowBuilds in your profile’s pnpm-workspace.yaml, and the install works on the next run. Allowing a build is a trust decision: only install sources you trust, and pin a commit (github:owner/repo#sha).

README

💡 What problem does it solve? Talking to an agent shouldn't mean only typing. BetterInput is an input-enhancement suite: prompt optimization, on-demand prompt templates, more local file formats you can bring into the input, and file-to-Markdown — plus the small UX refinements — making every input you feed an agent better (voice recognition is just one part of it).


🎬 Feature Demo

https://github.com/user-attachments/assets/caae08fc-2d8e-43c6-8bab-ade2d278337f

V0.1.5 version demo of four core features

✨ Implemented today

🗺️ Next (directions for better input)

BetterInput is a complete input-enhancement suite: not just one kind of input, but making every input you feed an agent smoother and easier. Next we go in three directions:

Files → structured (format upgrades)

📷 Image input: natively supported by DSH since rc.8 — the DeepSeek API supports image input natively, so we no longer ship an image plugin.

🎙️ Voice input: DSH now ships voice input natively (a desktop lab plugin built on a local model, which must be downloaded). This plugin keeps the browser Web Speech route — no downloads, no keys — as the lightweight option for the Web UI. The two routes are complementary.

Turn docs, sheets, and decks into clean, structured Markdown so the agent reads them at a glance.

  • 🧾 PDF → structured — PDF into an AI-friendly readable format (Markdown / plain text)
  • 📄 Office parsing — DOCX / PPT / XLSX into clean Markdown structure in one click
  • 🎬 Audio/video transcription — paste a local media file and get text (an upgrade to voice input)

Text & prompts

  • ✨ Prompt optimization — a one-click icon beside the input to have the AI polish / improve the prompt you wrote
  • 📝 Prompt template library — type / in the composer to search & insert common templates (coding / summarize / translate / role-play…)
  • 🧹 Text cleaning — paste messy / line-numbered / timestamped text and get clean copy
  • 🔤 Instant translation — one click to turn Chinese into English (or vice versa)
  • 📋 Smart paste — detect code / table / URL / quote on paste and wrap it appropriately

Interaction refinements

Input isn't just about features — it's also how comfortable and polished it feels.

  • 🎚️ Effort slider — removed; for slider-style reasoning-effort adjustment, install @HanaAyane/dsh-reasoning-effort
  • ✍️ Auto-complete suggestions — contextual continuations while you type, adopt in one click
  • 🧮 Variable fill — {{date}}, {{cwd}} and other tokens replaced automatically in the input

Planned around the directions; iterating continuously. Ideas welcome — file an Issue or open a PR.

💭 On settings toggles: for features that modify DSH's built-in plugins, adding another toggle in the settings screen is really redundant — if a feature can in principle be split out as its own plugin, it should be enabled/disabled by installing / uninstalling it, not by another switch inside BetterInput. So this plugin no longer pays the price of core-feature toggles: anything that can be split out already has been split out of BetterInput — install the standalone plugin when you want it, uninstall it when you don't.

🚀 Install

Prereqs: DeepSeek Harness (>= 0.2.0-rc.2) + Node.js ^22.19.0 || >=24.0.0 + Chrome/Edge.

💡 Pick either way. If you have the dsh CLI installed, use the short commands below. If not — or you don't want to install anything globally — use the npx full form: no global configuration needed at all. Published on npm.

Option A: global dsh CLI

# Install from npm (recommended)
dsh plugin --profile web add dsh-better-input

# Or from the GitHub repo
dsh plugin --profile web add github:DIAG5/dsh-better-input

# Uninstall
dsh plugin --profile web remove dsh-better-input

Option B: no dsh, or avoid global installs (npx full form)

Run dsh via npx — pulled on demand, nothing written to your global environment:

# Install from npm (recommended)
npx -y @deepseek-ai/dsh plugin --profile web add dsh-better-input

# Or from the GitHub repo
npx -y @deepseek-ai/dsh plugin --profile web add github:DIAG5/dsh-better-input

# Uninstall
npx -y @deepseek-ai/dsh plugin --profile web remove dsh-better-input

-y auto-confirms the download; the first run fetches the dsh CLI, cached by npx afterwards.

From source (development)

git clone https://github.com/DIAG5/dsh-better-input.git
cd dsh-better-input
npm install
npm run build
# with a global CLI:
dsh plugin --profile web add "$PWD"
# without a global CLI:
npx -y @deepseek-ai/dsh plugin --profile web add "$PWD"

Alternative: no package install — add a row to a preset's cordis.yml

If you already use an agent preset, just add one line (no install command needed):

- insert:
    - id: dsh-better-input
      name: dsh-better-input

After installing, refresh the Web UI — a microphone icon 🎤 appears on the right of the composer.

📖 Usage

1. Voice input

  1. Open any conversation and click the microphone button on the right of the input row.
  2. Start speaking — recognized text streams into the input in real time.
  3. Click again (or press Stop on the recognition bar) to finish.
  4. Review, edit, and send.

Recognition runs fully in your browser via the Web Speech API — no API key, no server round-trip. Auto-disabled in unsupported browsers (Firefox/Safari).

2. AI polishing

Settings → BetterInput → enable AI polishing → pick a model already configured in dsh.

The built-in prompt removes fillers, fixes homophone errors, restores punctuation, and formats spoken enumerations into numbered lists. Leave blank to use the built-in prompt (expand via Show the built-in prompt), or paste a custom prompt (the output-contract guard is always appended, so it returns clean text rather than answering).

3. Prompt optimization

  1. Type your prompt in the composer.
  2. Click the ✨ Optimize icon at the top right of the input row.
  3. After a short wait, a before / after comparison panel appears.
  4. Click Adopt to replace the draft with the optimized result, or Cancel to keep the original.

Thinking is off by default for fast, low-cost output. You can raise the effort tier in settings for deeper optimization.

4. Prompt templates

Save frequently used prompts (coding / summarizing / translating / role-play…) as templates and insert them on demand:

  1. Go to Settings → BetterInput → the "Prompt templates" section and click "New template"
  2. Fill in the name, description (optional), body, and tags (optional, comma-separated, used for search), then save
  3. Back in the composer, type / — template candidates pop up; keep typing to filter live by name / description / tag
  4. Pick one and the template body is inserted into the input, ready to edit further before sending

Templates are stored locally on the host machine at ~/.dsh/better-input/templates.json — nothing leaves your machine; up to 200 templates, 8,000 chars per body, sorted by most recently updated.

5. Add file / file-to-Markdown

  1. Click the 📎 Add file button at the top right of the composer to expand the file panel (click again to collapse).
  2. Click "Add file" to pick files (multiple allowed); they appear as small tags in the panel.
  3. Plain-text files (.txt / .md / .json / .py etc.) are marked ✓ immediately — type @ to insert and send them directly, no conversion. This is the "more file formats input" capability.
  4. Document files (.pdf / .docx / .xlsx etc.): click "Start conversion" — this is the "file-to-Markdown" capability:
    • Once converted they are marked ✓.
    • Type @ in the composer and pick the file from the candidates — it inserts as an @<filename> reference chip.
    • On send the chip expands into the converted Markdown body.
    • The panel keeps an "Edit" button so you can revise the converted result.
  5. Remove an unwanted file with ×.

Conversion runs locally via built-in parsers (PDF / Word / Excel / PPT / EPUB / HTML / CSV / JSON / XML etc.); the generated Markdown is sent with your message so the agent can read the document at a glance.

6. OCR vision recognition (scanned PDF / PPT)

For documents without a text layer — scanned PDFs, image-only PDFs, or PPTs whose slides are just pictures — regular conversion yields little or no text. Use OCR to let a vision model "read the pixels":

  1. First, in Settings → BetterInput, pick an "OCR vision model" (a vision model supporting image input; independent of the polish model)
  2. Add a .pdf / .pptx file and click "Start conversion" — you'll be asked "Use OCR?"
  3. Choose "Use OCR": PDF pages are rendered to images (or PPT embedded images extracted) and fed one at a time to the vision model as Markdown
  4. Choosing "Regular conversion" keeps the built-in text-layer extraction

If no OCR model is set, clicking "Use OCR" shows a neutral toast guiding you to Settings instead of a red error; if the chosen model explicitly declares it does not support image input, you're told upfront to switch, avoiding a confusing empty result.

7. Check for updates

  1. Open Settings → BetterInput → scroll to the bottom to the "About & Updates" section
  2. Click "Check for updates"
  3. If a newer version exists, it shows installed → latest plus an update command
  4. Run one of the commands below, depending on how you installed DSH:
    • With a global dsh CLI:
      dsh plugin --profile web update dsh-better-input
      
    • Without one, via npx:
      npx -y @deepseek-ai/dsh plugin --profile web update dsh-better-input
      

Note: DSH does not auto-update third-party plugins on launch — run the command above to pull the new release. This section simply helps you notice and follow updates promptly.

8. Settings

Setting Meaning
UI language Plugin copy supports Chinese / English, follows DSH's UI language switch, applied instantly
Recognition language Empty follows the browser language (e.g. zh-CN, en-US)
Recording limit 1–600 seconds, default 120, auto-stop
AI polishing On/off; when on, the transcript is auto-polished into the draft
Polish model A dsh model route
Polish reasoning effort Default: thinking off; optional higher tiers the model supports
Custom polish prompt Optional replacement of the built-in prompt
Prompt optimization On/off; when on, the ✨ button shows in the composer
Optimize model A dsh model route
Optimize reasoning effort Default: thinking off; optional higher tiers the model supports
Custom optimize prompt Optional replacement of the built-in optimize prompt
OCR vision model Vision model for scanned pages / embedded images; independent of the polish model. OCR is unavailable without one
Prompt templates Create / edit / delete templates (name, description, body, tags) in Settings; data is stored in a local JSON file on the host
About & Updates Shows installed version / license / repo, and a one-click "Check for updates" for the latest release and update command

Polish and optimization are configured independently — model, effort, and prompt each.

🧩 Compatibility

  • DeepSeek Harness >= 0.2.0-rc.2 (verified against 0.2.0-rc.2)
  • Node.js ^22.19.0 || >=24.0.0
  • Chromium-based browsers (Chrome / Edge)

🛠️ Development

npm install
npm run check    # typecheck
npm run build    # build lib/ (host ESM + browser bundle)

Client-only UI: npm run dev:watch, then refresh the UI. Host changes: restart dsh web.

🏗️ Architecture

  • src/index.ts — Host plugin entry, mounts the polish service
  • src/polish/service.ts — BetterInputPolishService (Typert remote): settings, dsh route discovery, LLM polishing & prompt optimization, file-to-Markdown (convertFile, via ctx.llm), prompt-template storage (templatesList / templatesSave / templatesRemove)
  • src/templates/ — prompt-template data model and host-side JSON store (atomic writes, corruption self-healing)
  • src/converter/ — pure-TypeScript file→Markdown conversion layer (PDF / DOCX / XLSX / PPT / EPUB / HTML / CSV / JSON / XML / ZIP), bundled on the Host only
  • src/about.ts — plugin identity and npm version check (About & Updates)
  • src/client/ — browser half: microphone/optimize/choose-file buttons (conversation.input.right), recognition bar/file panel (conversation.input.dock), settings page & template management (settings.section), @ reference-chip source (conversion-source), / template-candidate source (input trigger)
  • src/typert.ts / src/remote.ts — Client↔Host typed contract

📄 License

MIT


⭐ Support

This plugin is growing into a complete input-enhancement suite — think it's worth watching?

  • Give it a Star ⭐ (your stargazes fuel continued iteration)
  • File an Issue / open a PR
  • Share it with fellow DSH users

Thanks for your support ❤️

Content from the project README on GitHub ↗

Links

More in this category

View the whole category →

Community comments

Comments are public GitHub Discussions. Loading them connects to GitHub and Giscus; a GitHub account is required to post.