DeepSeek Harness Plugin

InkshadeWoods/dsh-tool-visual-primitives

Stars ★ 2 Downloads (30d) 1,359 Category Vision & Multimodal Added 2026-08-27 npm dsh-tool-visual-primitives

Analyzes conversation images through mode-specific prompts (caption, UI, document, grounding, topology, etc.) and injects structured evidence with coordinate primitives (boxes, points, refs) as text, with session-level caching for reuse across replay and compaction.

Install

# from npm (prebuilt)

dsh plugin --profile web add dsh-tool-visual-primitives

# from GitHub (first run asks for allowBuilds approval — follow the hint, retry)

dsh plugin --profile web add github:InkshadeWoods/dsh-tool-visual-primitives

Any plugin you install runs third-party code with your own permissions — it can read your files, use your credentials, and reach the network, and tool approvals don’t sandbox it. GitHub-sourced plugins also run build scripts at install time — pnpm blocks those until you allow them, so an install can stop with ERR_PNPM_GIT_DEP_PREPARE_NOT_ALLOWED or ERR_PNPM_IGNORED_BUILDS; dsh prints the exact key to add under allowBuilds in your profile’s pnpm-workspace.yaml, and the install works on the next run. Allowing a build is a trust decision: only install sources you trust, and pin a commit (github:owner/repo#sha).

README

A DeepSeek Harness (DSH) plugin that gives text-only chat models a visual capability. It sends an image to an external vision model, then returns plain-text visual evidence—optionally with spatial references—to the original chat model. The chat model itself does not need native image input.

The design is inspired by DeepSeek's Thinking with Visual Primitives: normalized coordinates and referenceable objects turn image understanding into evidence that later reasoning can inspect and use.

The package is published on npm. The official DSH one-command install below is recommended; local source mounting remains available for development and local debugging.

What's New in 1.5.3

  • Declares compatibility with DSH 0.2.0-rc.2: verified end-to-end in real Web and desktop sessions (conversation-enhancement bridging, the vision_analyze tool, credential reads/writes, and the settings panel all work). Since 0.2.0 the version gate checks both engines.dsh and peerDependencies; this plugin's open-ended >=0.1.2-rc.1 floors on both sides pass without any exemption.
  • Fixed Invalid schema for function 'vision_analyze' when chatting through official models: the tool's parameter declaration is now a standard JSON Schema with an object root, matching what the host's defineTool compiles to. The host projects tool parameters verbatim to the model API, and DeepSeek official endpoints strictly require an object-rooted schema — the previous property-map shape was rejected. Custom models are unaffected.
  • Fixed 403 on the desktop settings panel: the desktop front end loads through an Electron custom protocol without a Referer, so the same-origin check now allows loopback requests that lack a source header or carry a non-http/https protocol origin; malicious web pages always send an http/https Origin and cannot enter the new branches, so security is not weakened.
  • Requirements updated: DSH 0.1.2-rc.1 or newer is required, verified up to 0.2.0-rc.2.

What's New in 1.5.2

  • Declares compatibility with DSH 0.1.7-rc.2: the host contract surface between 0.1.5-rc.2 and 0.1.7-rc.2 was fully compared (ctx.llm adapter and routing, ctx.credentials, ctx.attachments.readImage, ctx.fs, the ctx.remote.credentials RPC, webServer route registration, and all four client injection points are compatible) and verified in a real session.
  • Fixes two SSRF-guard defects (vision/image-source.mjs): ① the ::ffff:0:0/96 block rule was removed — Node's BlockList stores every IPv4 address internally in IPv4-mapped form, so that rule rejected every public IPv4 host (the IPv4 rules already cover the mapped form); ② the net.connect custom lookup callback now uses the all: true array contract and reports an explicit error when every resolved address is blocked — previously a single-address string was silently dropped by Node and an empty result passed undefined, both surfacing as ERR_INVALID_IP_ADDRESS.
  • Attachment read error codes now carry the original message (vision/analysis-core.mjs): non-ATTACHMENT_* failures (e.g. the plain Error raised by the SSRF guard) are no longer collapsed into a bare ATTACHMENT_READ_ERROR; the stable suffix stays for machine matching while the real cause becomes diagnosable.
  • Notes on the new 0.1.7 host capabilities: the host adds an IMAGE_OFFLOAD_REQUIRED image-budget mechanism and text-only image projection helpers, but the bridge model declares image input, so the host forwards images to the bridge adapter unchanged, where the plugin converts them into visual-evidence text before forwarding to the underlying model — no conflict with the host mechanisms.
  • GUI model selection in 0.1.7 filters by catalog membership: the plugin's listModels dynamically lists every enabled route's bridge model, so configured bridges appear and can be selected normally.

What's New in 1.5.1

  • Declares compatibility with DSH 0.1.5-rc.2: verified end-to-end in a real session (conversation-enhancement bridge calls, credential read/write, settings and model-catalog routes, and attachment reads all work).
  • Adds an explicit @deepseek-ai/dsh peer dependency (>=0.1.2-rc.1), matching the existing engine declaration, so npm install resolution checks the host version.

What's New in 1.5.0

  • Fixed "DSH credential service unavailable" when saving settings on DSH 0.1.2-rc.1: credential reads/writes moved to the new ctx.remote.credentials API, with a compatibility path for older hosts.
  • The "Conversation enhancement" model picker now greys out models that already support image input natively (they need no bridging) and labels them accordingly.
  • Gateway image rejections (e.g. "Model do not support image input") and empty vision-model responses now produce clear error messages instead of injecting raw gateway JSON into the conversation.
  • Fixed persisted settings being overwritten by null values and unbounded model-catalog response reads.
  • Stability polish: retry-path diagnostics, empty-response interception, and concurrency-dedup documentation.

What's New in 1.4.0

  • The "Chat visual models" list in settings now reads through the new model-catalog API of DSH 0.1.2-rc.1, matching the new host.
  • The minimum supported DSH version is raised to 0.1.2-rc.1 (engine declarations in dsh.plugin.json and package.json tightened accordingly).
  • Client injection dependencies trimmed: @deepseek-ai/dsh-client-runtime is no longer injected.

What's New in 1.3.0

  • Compatible with the new DSH host adapter interface (0.1.1-rc.8 and above).
  • Auxiliary session calls (title generation, context compaction) no longer trigger duplicate visual analyses, reducing quota usage and occasional failures.
  • Concurrent analyses of the same image and prompt are merged into a single upstream request.
  • New "Diagnostics log" setting (off by default) to output plugin runtime logs for troubleshooting.
  • Connection test improvements: uses the real token budget, and failures include an upstream response summary in the error message.
  • Install commands updated to the direct dsh form (requires DSH installed globally).

What's New in 1.2.0

  • Reworked the settings hierarchy and connection flow: completed connections collapse by default, with full editing and Save & Test Connection available when expanded.
  • Improved vision-model selection, model discovery, custom model IDs, connection status, and error feedback.
  • Optimized multi-turn image conversations: text-only follow-ups no longer repeat vision requests, while explicit references to earlier images can reuse cached visual evidence.
  • Added a clear, provider-neutral error message for image-content protocol incompatibility.

Highlights

  • Two entry points backed by one vision_analyze core: an explicit tool and [vision] chat-model variants.
  • Automatic detection of 11 visual tasks: captioning, inventory, multi-subject, counting, grounding, spatial relation, comparison, path tracing, topology, UI, and document visual analysis.
  • Three output-detail levels: brief, standard (default), and verbose.
  • Three visual-primitive policies: auto (default), on, and off. Primitives use <ref>, <box>, and <point> with normalized 0–999 coordinates.
  • Appends [vision] variants only for the text-only chat models you choose; original models stay unchanged.
  • Session-scoped evidence caching: a follow-up reuses evidence only when it covers the new question; otherwise the image is read again.
  • Native settings page for secure credential storage, connection checks, searchable /models discovery, custom model IDs, and collapsible provider/model selection (native image models are greyed out).

How It Works

Image + user question
        │
        ▼
detectVisionMode() → shouldUsePrimitives() → buildVisionPrompt()
        │
        ▼
External vision model (OpenAI-compatible Chat Completions)
        │
        ▼
Plain-text visual evidence (optionally with primitives)
        │
        ▼
Original text model continues the answer

Mode and Detail are orthogonal: Mode selects the task, while Detail controls information density. Primitives decides whether structured spatial evidence is required.

Requirements

  • A working DSH Web Profile (DSH 0.1.2-rc.1 or newer is required, verified up to 0.2.0-rc.2).
  • Node.js >= 20 and pnpm.
  • An accessible vision-model service. By default the plugin uses OpenAI-compatible endpoints:
    • POST <Base URL>/chat/completions
    • Optional model discovery: GET <Base URL>/models

The vision provider and the enhanced text-model provider may be different.

Installation

One-command npm install

Run this in PowerShell or a terminal (the dsh command below assumes DSH is installed globally via npm install -g @deepseek-ai/dsh; without a global install, replace dsh with npx @deepseek-ai/dsh):

dsh plugin --profile web add dsh-tool-visual-primitives@latest

The command invokes DSH's official plugin manager and installs the package into the web Profile. After installation, DSH detects this package's dsh.bundle.patch declaration, adds the plugin to dsh.profile.bundles, and applies the package's own cordis.patch.yml during boot.

It is safe to run repeatedly: it does not duplicate the bundle and does not directly change the Profile's own cordis.patch.yml or other plugin settings, so it can coexist with plugins such as dsh-better-sidebar.

If you use a Profile other than web, replace --profile with that Profile name:

dsh plugin --profile <your-profile-name> add dsh-tool-visual-primitives@latest

Restart DSH completely after installation, then configure API Key, Base URL, and the vision model under Settings → Visual Analysis.

Install from a Local GitHub Clone

This is the installation path verified for the current release. You may use another source directory if you also update $source.

$source = 'D:\DSH\dsh-tool-visual-primitives'
git clone https://github.com/InkshadeWoods/dsh-tool-visual-primitives.git $source

Set-Location $source
pnpm install
pnpm run build

dsh plugin --profile web add $source

The official DSH CLI adds the local package to the Profile dependencies and maintains dsh.profile.bundles from the package's dsh.bundle.patch declaration. No manual edit of the Profile's package.json or cordis.patch.yml is needed.

Restart DSH completely:

dsh web

When the client UI changes for the first time, force-refresh the browser with Ctrl+Shift+R.

Uninstall

If you also want to remove the saved API key, select Clear API Key in the plugin settings first. Then run DSH's official uninstall command:

dsh plugin --profile web remove dsh-tool-visual-primitives

DSH removes the package dependency and automatically removes its bundle from dsh.profile.bundles. Restart DSH when it completes.

First-Time Setup

Open Settings → Vision Analysis in DSH, then:

  1. Enter the API Key, Base URL, and vision model.
  2. Select Load models to retrieve <Base URL>/models; the list is searchable.
  3. If the service does not expose a model directory, enter a Custom model ID instead.
  4. Select Test connection.
  5. Configure analysis parameters, then choose the text-only chat models that should receive a [vision] variant.

Saved API keys are never shown again when the page is reopened. Entering a new value replaces the old key; Clear API Key removes it.

API and Model Discovery

Item Behavior
Vision request POST <Base URL>/chat/completions with Authorization: Bearer <API Key>
Model discovery GET <Base URL>/models with Accept: application/json and the same API key
Discovery failure A custom model ID remains available and does not prevent vision analysis
Xiaomi Mimo URLs The plugin automatically uses the api-key header

Analysis Parameters

Setting Options / default Purpose
Visual primitives auto / on / off (default: auto) auto decides from Mode and Detail; on forces coordinate-based evidence; off requests plain-text evidence only.
Analysis detail brief / standard / verbose (default: standard) Controls output density; it does not alter task detection.
Retry mode off / on / format-only (default: off) When required primitives are missing, on re-reads the image; format-only preserves conclusions where possible and repairs formatting.
Maximum image size 10 MB Applies to local files, remote images, and chat attachments.
Timeout 180000 ms Maximum wait for one vision-model request.
Output-token budget auto or manual (default: auto) auto follows Detail: brief 1024, standard 2048, verbose 4096.
Diagnostics log off / on (default: off) When enabled, plugin runtime logs are printed to the DSH console for troubleshooting.

The 11 Auto-Detected Modes

Mode Best for Evidence focus
caption "What is this image?" Overall summary and key objects
object_inventory "What objects are present?" Main-object list and positions
multi_subject "Who is shown left to right?" Subject ordering, features, and positions
counting "How many buttons?" Candidates, exclusions, and count
grounding "Where is the red button?" Target and candidate locations
spatial_relation "Which side is A on?" Relative position, occlusion, containment
comparison "Compare these two areas" Dimensions and visible evidence for each side
path_tracing "How does the route go?" Start, key points, end, uncertainty
topology "Is the maze solvable?" Connectivity, blockers, and conclusion
ui_analysis "How do I use this screen?" UI elements, state, location, next step
document_visual "Explain this chart/poster" Headings, text blocks, tables, reading order

The highest-priority keyword match wins; caption is used when nothing matches. For example, "How many buttons are on this screen?" selects ui_analysis and then applies the selected Detail level.

Usage

Option 1: Chat with a [vision] Model

  1. In Chat visual models settings, enable a text-only model.
  2. Reopen the model list and select the new Model name [vision] entry.
  3. Upload, paste, or drop an image, then ask your question normally.

The plugin replaces only image blocks with visual-evidence text. The selected original text model still produces the final answer.

For a follow-up that explicitly refers to the latest image, the plugin checks whether cached evidence covers the new task, detail level, and requested objects/locations. It re-analyzes the image when coverage is insufficient rather than treating an incomplete previous answer as fact.

Option 2: Explicit vision_analyze Tool

The tool accepts exactly one image source: a local absolute path or an HTTP(S) URL.

{
  "image_path": "D:/images/dashboard.png",
  "prompt": "Count the main clickable buttons and identify their positions"
}
{
  "url": "https://example.com/chart.png",
  "prompt": "Explain the chart trend and identify unreadable labels"
}

Remote URLs cannot target localhost, private-network addresses, or include credentials. Redirects are rejected to reduce server-side request forgery risk.

Evidence Format

When primitives are enabled, the external vision model is instructed to return evidence under these headings:

[Mode]
[Visual Primitives]
[Observations]
[Relations]
[Uncertainty]
[Answer]

Location evidence looks like this:

<ref>submit_button</ref><box>[[742, 861, 900, 930]]</box>
<point>[[125, 430], [210, 430], [300, 510]]</point>

All coordinates are relative 0–999 values, not source-image pixels.

Verified End-to-End Scenarios

All assets and results are in test/.

Scenario Verified outcome
Chat image understanding A [vision] model successfully read a DSH usage-mode comparison image and supplied a structured description to the text model.
Screenshot-driven UI recreation A [vision] model interpreted a Bilibili home-page screenshot; the text model then generated an independent Bilibili-style HTML page from that evidence.

Image-understanding result

Successful image reading through the chat-vision entry

UI recreation flow and result

Conversation that recreates a UI from a screenshot

HTML page generated from visual evidence

These outcomes validate the current end-to-end chain. Generated results still depend on the external vision model, text model, prompt, and image quality.

Development

pnpm install
pnpm run build

pnpm run build generates the client bundle at lib/client.js.

  • The server entry is index.mjs; restart DSH after changing it.
  • After changing the client UI, rebuild the package and force-refresh the browser page.

License

MIT

Acknowledgements and References

Content from the project README on GitHub ↗

Links

More in this category

View the whole category →

Community comments

Comments are public GitHub Discussions. Loading them connects to GitHub and Giscus; a GitHub account is required to post.