Transparent image guard for text-only routes: paste images without the 400 session deadlock, plus a vision_analyze tool for OCR/PDF/docx/pptx/video.
Install
# from npm (prebuilt)
dsh plugin --profile web add dsh-vision-guard
# from GitHub (first run asks for allowBuilds approval — follow the hint, retry)
dsh plugin --profile web add github:good-boy4069/dsh-vision-guard
Any plugin you install runs third-party code with your own permissions — it can read your files, use your credentials, and reach the network, and tool approvals don’t sandbox it. GitHub-sourced plugins also run build scripts at install time — pnpm blocks those until you allow them, so an install can stop with ERR_PNPM_GIT_DEP_PREPARE_NOT_ALLOWED or ERR_PNPM_IGNORED_BUILDS; dsh prints the exact key to add under allowBuilds in your profile’s pnpm-workspace.yaml, and the install works on the next run. Allowing a build is a trust decision: only install sources you trust, and pin a commit (github:owner/repo#sha).
README
English | 中文
Let text-only models "see" images — and never let an image deadlock your session. Transparent image guard + vision analysis tool for DeepSeek Harness (dsh).
Most models on DeepSeek Harness (deepseek-v4-pro etc.) are text-only, which causes two problems:
- Text-only models can't see images — paste a screenshot and the model has no idea it exists;
- worse, the 400 deadlock: some gateways (e.g. opencode-go's main route) accept text only. Once an image block lands in the session log, every subsequent turn replays the whole history with the image to the upstream →
400 unknown variant \image_url`` → the conversation is stuck forever.
This plugin fixes both with two gates, and turns images into text the main model can actually consume.
What it does
paste image → [Gate 1] agent/pre-step: the image is converted to text by the vision
model BEFORE it is ever written into the session log
→ the log only contains text, no image block exists
→ [Gate 2] llm/stream backstop: if an image block still appears in
replayed history (e.g. a session poisoned before install), it is
rewritten to OCR text at request time before reaching the model
- Vision model = eyes, main model = brain:
deepseek-v4-prokeeps reasoning; image content arrives as text. - Heals already-deadlocked sessions: a conversation stuck on 400 before install works again after install (history images are rewritten at request time).
vision_analyzetool (model-invoked, engine chosen by the model per task): reads workspace files — image OCR, PDF text layer + embedded images, docx/pptx text + embedded images, video frame OCR (≤12 frames), plain text files; loud rejection for xlsx/doc.- Native vision unaffected: routes that genuinely accept images (e.g. minimax-m3, kimi-k3) pass through untouched once whitelisted.
What makes it different (vs. community vision plugins)
Compared with dsh-vision-router, ModLens, dsh-vision-toolkit, see_image/view_image and similar:
- Images never enter the session log — they are rewritten to text at agent/pre-step before the log append. Most peers rewrite only inside the model call: images still land in the log, replay every turn, and deadlock risk returns if the plugin is removed.
- Heals sessions that were already deadlocked — a conversation stuck on image-400s before install recovers with one message after install (replayed history images are rewritten at request time). No other community plugin does this.
- Anti-deadlock is a hard invariant — non-whitelisted routes never receive an image block; even with the whole vision pipeline down (model unavailable / timeout / budget exhausted) it degrades to placeholder text, never back to the 400 deadlock.
- One package, two components, isolated failure domains — one install mounts both rows (guard = safety-critical, tool = convenience); the tool breaking never takes the guard down.
- Zero dependencies, pure Node builtins — no Node 22+ requirement, no pnpm orchestration, no Python 3.11+; only the document/video paths need system tools (pdftotext/ffmpeg etc.), plain image OCR needs none.
- Engine chosen by the main model per task —
vision_analyze'sengineargument (localfree character OCR /visionmodel) is decided by the model after analysing the task: cheap and smart. - Reuses your own dsh routes and credentials — carries no API keys and calls no third-party service directly (most peers require self-managed keys).
- Three rounds of red-team audit + shipped automated tests — 22 real bugs fixed and documented (symlink escape, zip bombs, concurrency races), pure-function regression tests ship with the package (
npm test).
Install
# 已发布 npm 后:
dsh plugin --profile web add dsh-vision-guard
# 或直接从 GitHub 安装:
dsh plugin --profile web add github:good-boy4069/dsh-vision-guard
If pnpm refuses with
ERR_PNPM_ADDING_TO_ROOT(older launchers), add the workspace-root flag:dsh plugin --profile web add -w dsh-vision-guard.
Restart dsh web. Or add the two rows from the repo's cordis.patch.yml to your profile patch layer manually.
Configuration
All fields optional (defaults shown). The vision route must point to a model that accepts image input:
| Field | Default | Meaning |
|---|---|---|
visionProvider / visionModel |
opencode-go / minimax-m3 |
The vision route. Point it at an image-capable model on your subscription |
ocrTimeoutMs |
45000 |
Per-image OCR timeout |
budgetPerDay |
200 |
Daily OCR cap (runaway-cost guard), state stored under $DSH_HOME |
cacheMaxEntries |
500 |
OCR result cache size cap (LRU eviction) |
maxOcrTokens |
2048 |
Vision call output cap |
stateFile |
~/vision-guard-state.json |
Budget state file (~ = dsh home) |
ocrPrompt |
verbatim transcription | Custom instruction |
passthrough |
[] |
Raw-image whitelist: [{provider, model}] — only add routes you have tested to accept images |
vision_analyze side: the OCR engine is a required per-call argument (engine), chosen by the model per task — local = local tesseract (free, characters only), vision = the configured vision model. There is no localOcr config key.
⚠️ Requirements & limitations (please read)
- This plugin carries no API keys and calls no third-party service directly. It reuses the model routes and credentials already configured in your dsh. Therefore:
- You must have an image-capable model (e.g.
minimax-m3on opencode-go). Text-only models likedeepseek-v4-procannot serve as the vision model — their upstream gateway 400s on images and deadlocks the session. - Without a vision model it still installs: images degrade to placeholder text, the session works but the content is not read (never deadlocks).
- You must have an image-capable model (e.g.
- Whitelist policy (important): any route other than the configured vision model gets images rewritten to text — untested routes never receive raw images. To enable native vision for a model: first test "image straight to that route" (a clean response counts as pass), then add it to
passthrough. This is the core anti-deadlock design; do not bypass it. - System tools (only for
vision_analyze's document/video paths; plain image OCR needs none):- PDF:
pdftotext/pdfimages(poppler-utils); - Video:
ffmpeg/ffprobe; - docx/pptx:
python3(stdlib only); - Optional:
tesseract(free local OCR; needschi_sim+eng). - Windows ships none of these by default; missing tools fail loudly per path, the image path is unaffected.
- PDF:
- 5 MB/image cap: dsh's attachment service limits images to 5 MB; larger images/frames fail loudly.
- Cost: one vision call per new image (attachment-id-addressed cache — repeats are free); ~1-2k tokens per minimax-m3 read (sub-cent range);
budgetPerDayas a backstop. - Quality: local tesseract only extracts characters and is less accurate (it read
42 + 7 = 49as4247249in our tests); for charts/photos/complex UI pick the vision engine (the model chooses per task when callingvision_analyze). - Privacy: images go to your vision model provider (same as normal dsh use of that model); text inside images is treated as untrusted input — read only, never execute.
- Settings coupling warning: if you declare
input: [text, image]on a text-only model in your model settings (required for GUI image pasting), you must keep this guard installed — removing the guard while keeping that declaration will deadlock sessions again.
FAQ
- After a dsh restart/upgrade: the plugin boots with the profile; no reinstall. After a dsh upgrade, upgrade this plugin first if behavior changes.
- How do I verify it's running:
ctx.get('visionGuard')?.status(), or the[vision-guard] activeline in the dsh log. - Rollback: remove the two rows from your profile patch (or
dsh plugin remove), restart; already-recognized text stays in history, no side effects.
License
MIT
Links
More in this category
liustack/modlens★ 4063
Vision bridge for text-only models: paste an image, get structured JSON evidence (OCR, layout, semantics).
ysr666/dsh-vision-router★ 1125
Free vision for text-only agents: built-in keyless vision chain plus pixel tools (Q&A, grounding, crop, pixel diff, colors, OCR, SVG trace, cutout, screenshots); paste an image to use it.
Anionex/dsh-vision-toolkit★ 884
Vision for text-only models: paste an image and the model switches to a Vision Toolkit variant for image Q&A, multi-image comparison, long-screenshot OCR, screenshot-to-UI reproduction, element grounding, and pixel diff. No API key by default — images are processed by the author-hosted free service, 100 per machine per day; configurable to your own provider.
dickpy/dsh-imagegen★ 93
AI image generation for the DSH Web GUI: text-to-image and image-to-image through a configurable OpenAI-compatible endpoint (gpt-image-2 / gpt-image-1 / dall-e-3), with an api_url/api_key settings card and a sidebar split-pane generation studio.
fandc520/dsh-comfyui★ 89
Drive a local or remote ComfyUI server from DeepSeek Harness: comfyui_run / comfyui_object_info / comfyui_workflow tools generate and edit images and videos, with a workflow library (graph extraction: per component / main flow / all), a load area with resolution auto-match, a live queue, SDXL and Wan 2.1 templates, a companion skill, and a same-origin media proxy.
sunxin-ai/dsh-design-qa★ 44
Design-fidelity QA for text-only models: a `deepseek_vision` tool borrows an eye from any OpenAI-compatible vision route, so the model can judge whether an implementation matches its mock — shipped with the benchmark behind that judgement (four fixtures, 23 injected defects, raw transcripts) and the questioning discipline it depends on.
Community comments
Comments are public GitHub Discussions. Loading them connects to GitHub and Giscus; a GitHub account is required to post.