Enhanced vision toolbox: 14 pixel-level vision tools (describe, ground, detect, crop, pixel-diff, OCR, long-screenshot OCR, vectorize, colors, cutout, screenshot, present, materialize, html-screenshot) driven by one OpenAI-compatible endpoint, with clean \[图片: path] bridge markers, content-safety classification and rate-limit auto-retry.
Install
# from GitHub (first run asks for allowBuilds approval — follow the hint, retry)
dsh plugin --profile web add github:xing666173/dsh-vision-hub#path:/tool-vision
Any plugin you install runs third-party code with your own permissions — it can read your files, use your credentials, and reach the network, and tool approvals don’t sandbox it. GitHub-sourced plugins also run build scripts at install time — pnpm blocks those until you allow them, so an install can stop with ERR_PNPM_GIT_DEP_PREPARE_NOT_ALLOWED or ERR_PNPM_IGNORED_BUILDS; dsh prints the exact key to add under allowBuilds in your profile’s pnpm-workspace.yaml, and the install works on the next run. Allowing a build is a trust decision: only install sources you trust, and pin a commit (github:owner/repo#sha).
README
GitHub: Scorp1o117/dsh-tool-vision · npm: dsh-tool-vision
Part of the DeepSeek Harness Enhancement Suite — Vision · Soul/Persona · Long-term Memory · Plugin Marketplace.
External vision model for DeepSeek Harness.
DeepSeek's own models are text-only, and the harness derives every model
request strictly from the session log (llm/stream requests must equal the
durable derivation — the agent-loop invariant). This plugin bridges the gap
in two ways:
inspect_imagetool — sends an image (local file, or http(s) URL) to any OpenAI-compatible/chat/completionsendpoint that supportsimage_urlcontent parts, and returns the vision model's textual answer into the agent loop.- Image bridge (v0.2.1) — pasted images are turned into
inspect_imagehints before they enter the durable log, on theagent/pre-stepwaterfall (the one seam where the harness lets a plugin replace the messages of a proposed step). Images already logged by an older version are repaired lazily with a surfacereplaceon the session's first pre-step. Only models listed inmultimodalModelsreceive image blocks directly; a model's declaredinputModalitiesare never consulted, because profiles routinely declareinput: [text, image]on text-only models just to pass the harness's prompt-admission check.
- Zero dependencies beyond the dsh SDK — works with any compatible endpoint: OpenAI GPT-4o, Qwen-VL (DashScope), GLM-4V (Zhipu), Moonshot, Gemini compatible endpoints, local Ollama, etc.
- Registered on the global tools layer: every agent in the process can
call
inspect_image. - Web UI settings section (v0.3.0): Settings → 视觉模型 edits the
tool-visionnamespace (API endpoint, write-only key, model, bridge options) insettings.yaml; changes hot-apply without a restart. The API key lives insettings.yaml, not the profile patch. Mount by package name (name: 'dsh-tool-vision') so the web client bundle is discovered.
Install
Mount in a profile patch ($DSH_HOME/profiles/<name>/cordis.patch.yml):
- insert:
- id: tool-vision
name: 'dsh-tool-vision' # after: pnpm add dsh-tool-vision in the profile
config:
baseURL: 'https://api.openai.com/v1'
apiKeyEnv: 'VISION_API_KEY'
model: 'gpt-4o-mini'
Or load it from a local path without npm:
- id: tool-vision
name: './plugins/dsh-tool-vision/index.js'
Config
| Field | Default | Meaning |
|---|---|---|
baseURL |
https://api.openai.com/v1 |
OpenAI-compatible API base URL. |
apiKey |
'' |
API key (takes precedence over env). |
apiKeyEnv |
VISION_API_KEY |
Env var holding the key. |
model |
gpt-4o-mini |
Vision model id. |
maxTokens |
1024 |
Max output tokens. |
timeoutMs |
60000 |
Per-request timeout. |
maxImageBytes |
10MB |
Largest accepted local image. |
description |
default | Tool description shown to the model. |
bridgeTextOnly |
true |
Bridge pasted images to text hints on models that cannot see images. |
bridgeExportDir |
temp | Export dir for bridged images (os.tmpdir()/dsh-vision-bridge). |
multimodalModels |
[] |
Model ids that receive image blocks directly (e.g. mimo-v2.5). |
Image bridge setup
- In your model settings, declare image input on the models you paste
images onto, so the harness admits image messages (pi-ai style):
llm-pi-ai: providers: your-provider: models: - id: deepseek-v4-flash input: [text, image] - List genuinely multimodal models in the plugin config so they receive
image blocks untouched:
- id: tool-vision name: 'dsh-tool-vision' config: multimodalModels: ['mimo-v2.5', 'grok-4.5']
Then pasting an image while on a text-only model stores a hint like
[User sent an image, exported to: <path>. Inspect it with the inspect_image tool...]
in the transcript (the pasted image no longer renders as pixels in that
message), and the agent inspects it through the configured vision endpoint.
Why not
llm/stream? The harness freezes every request and the agent-loop invariant fails any request whose messages diverge from the session-log derivation (log-reconstruction desync), and this cordis waterfall'snext()cannot replace request arguments. Theagent/pre-stepwaterfall is the supported seam: its decision messages become the durable log, so the invariant stays satisfied.
Key resolution order: config.apiKey → process.env[apiKeyEnv] →
process.env.OPENAI_API_KEY.
Tool: inspect_image
| Arg | Required | Meaning |
|---|---|---|
path |
✅ | Image path (absolute, or relative to the current workspace) or http(s) URL. |
question |
– | Optional specific question about the image. |
detail |
– | auto / low / high resolution hint. |
Example endpoints (baseURL):
- OpenAI:
https://api.openai.com/v1—gpt-4o,gpt-4o-mini - Alibaba DashScope (Qwen-VL):
https://dashscope.aliyuncs.com/compatible-mode/v1—qwen-vl-plus,qwen-vl-max - Zhipu (GLM-4V):
https://open.bigmodel.cn/api/paas/v4—glm-4v-flash(free tier),glm-4v-plus - Moonshot (Kimi):
https://api.moonshot.cn/v1—moonshot-v1-8k-vision-preview - Ollama local:
http://localhost:11434/v1—llama3.2-vision(no key)
Note for users
- This plugin is a standard profile bundle (
dsh.bundle.patch):dsh plugin --profile web add dsh-tool-visioninstalls and mounts it in one step — no manualcordis.patch.ymledits needed.- The settings section needs the
dsh-host-apiproxynamespace allowlist; the plugin patches it automatically on first start — restartdsh webonce more and the section appears. A dsh update overwrites the patch; the next plugin start re-applies it.- Settings changes hot-apply (no restart needed).
- Tested against DSH
0.1.0-rc.6.
Pixel-level vision tools (v0.4.0, ported from dsh-vision-router)
14 vision_* tools driven by the same configured endpoint as
inspect_image (baseURL/apiKey/model) — no provider chain, no local models,
no extra settings:
| Tool | Purpose |
|---|---|
vision_describe |
Image Q&A / multi-image comparison (optional structured JSON) |
vision_ground |
Locate a target and return its ORIGINAL-pixel bounding box |
vision_detect |
Enumerate elements (buttons, inputs, icons…) with numbered boxes |
vision_crop |
Crop a pixel region to a PNG artifact |
vision_pixel_diff |
Per-pixel comparison: ratio, worst regions, heatmap, report |
vision_colors |
Dominant-color quantization for palette matching |
vision_ocr |
Verbatim text transcription (letters only — not scene analysis) |
vision_long_screenshot_ocr |
Chunked long-screenshot transcription into Markdown |
vision_trace |
Potrace vectorization into colored SVG (worker-thread, safe) |
vision_extract_foreground |
Solid-background removal → transparent PNG |
vision_html_screenshot |
Headless render of a local .html (network blocked) |
vision_screenshot |
Desktop capture (Win: PowerShell; macOS: screencapture; Linux: import/scrot) |
vision_present |
Publish a generated image to the user via the host attachment store |
vision_materialize |
Copy an attachment/local image into the workspace as a real path |
Quality & safety details:
- Content-hash cache keyed by endpoint+model+image+question (no stale answers across model switches, failures are never cached).
- Uniform 4MP downscale before every model call; oversized inputs are rejected with a clear error (stat pre-check, 20MB cap on both file and attachment paths).
- Rate-limit / 5xx auto-retry with Retry-After-aware backoff; endpoint
content-safety rejections are surfaced as
VISION_CONTENT_FILTEREDinstead of a generic backend error. - Long-OCR bounds: 120s total budget, 40-chunk cap, cancellation checks, stop-on-first-backend-failure.
- Path containment for relative inputs; artifacts land in
<workspace>/.dsh-tool-vision/. - Bridge marker: pasted images become a short
[图片: <path>]marker and the usage rule lives in a system-prompt section (clean transcript, same model behavior).
Requires sharp / potrace / puppeteer-core (declared in dependencies;
missing ones degrade lazily with an install hint and never break other tools).
Limitations
- A bridged image enters the conversation as a text hint (a transcript, not
pixels) — pixel-precise in-context reasoning is not available to text-only
models; the vision model's description comes back through
inspect_image. - Images are base64-transferred; mind privacy and size limits.
- Independent of the dsh-llm routing/retry system; failures return clear errors to the agent.
License
MIT
Links
More in this category
liustack/modlens★ 4142
Vision bridge for text-only models: paste an image, get structured JSON evidence (OCR, layout, semantics).
ysr666/dsh-vision-router★ 1132
Free vision for text-only agents: built-in keyless vision chain plus pixel tools (Q&A, grounding, crop, pixel diff, colors, OCR, SVG trace, cutout, screenshots); paste an image to use it.
Anionex/dsh-vision-toolkit★ 885
Vision for text-only models: paste an image and the model switches to a Vision Toolkit variant for image Q&A, multi-image comparison, long-screenshot OCR, screenshot-to-UI reproduction, element grounding, and pixel diff. No API key by default — images are processed by the author-hosted free service, 100 per machine per day; configurable to your own provider.
dickpy/dsh-imagegen★ 100
AI image generation for the DSH Web GUI: text-to-image and image-to-image through a configurable OpenAI-compatible endpoint (gpt-image-2 / gpt-image-1 / dall-e-3), with an api_url/api_key settings card and a sidebar split-pane generation studio.
fandc520/dsh-comfyui★ 98
Drive a local or remote ComfyUI server from DeepSeek Harness: comfyui_run / comfyui_object_info / comfyui_workflow tools generate and edit images and videos, with a workflow library (graph extraction: per component / main flow / all), a load area with resolution auto-match, a live queue, SDXL and Wan 2.1 templates, a companion skill, and a same-origin media proxy.
sunxin-ai/dsh-design-qa★ 44
Design-fidelity QA for text-only models: a `deepseek_vision` tool borrows an eye from any OpenAI-compatible vision route, so the model can judge whether an implementation matches its mock — shipped with the benchmark behind that judgement (four fixtures, 23 injected defects, raw transcripts) and the questioning discipline it depends on.
Community comments
Comments are public GitHub Discussions. Loading them connects to GitHub and Giscus; a GitHub account is required to post.