Routes images to a vision model of your choice - auto-rewrite, explicit tools, or hybrid - so a text-only chat model does not fail a turn that contains a picture.
Install
# from npm (prebuilt)
dsh plugin --profile web add @goodandready/dsh-vision-bridge
# from GitHub (first run asks for allowBuilds approval — follow the hint, retry)
dsh plugin --profile web add github:GooDAnDReaDY/dsh-vision-bridge
Any plugin you install runs third-party code with your own permissions — it can read your files, use your credentials, and reach the network, and tool approvals don’t sandbox it. GitHub-sourced plugins also run build scripts at install time — pnpm blocks those until you allow them, so an install can stop with ERR_PNPM_GIT_DEP_PREPARE_NOT_ALLOWED or ERR_PNPM_IGNORED_BUILDS; dsh prints the exact key to add under allowBuilds in your profile’s pnpm-workspace.yaml, and the install works on the next run. Allowing a build is a trust decision: only install sources you trust, and pin a commit (github:owner/repo#sha).
README
⚡ Overview & The Problem
When interacting with text-only chat models (any provider/model that accepts text only) in DeepSeek Harness, users cannot natively attach and send images:
- In DSH 0.1.2-alpha.2+, the core session controller performs a strict server-side modality check (
ctx.llm.resolveModelInfo). If the active conversation model lacksimageininputModalities, the prompt is immediately rejected with asession/attachment-invaliderror ("Model does not support image input"). - Standard text-only adapters throw errors when encountering raw multimodal image blocks in their message payload.
How dsh-vision-bridge Solves This
dsh-vision-bridge acts as an intelligent intermediary inside the Cordis runtime:
- Server-Side Modality Bridge (v0.5.3+): Decorates
ctx.llm.resolveModelInfoandctx.llm.listModelsso the session controller accepts image attachments on all models when bridging is active. - Automatic Image Rewrite (
agent/pre-step&llm/stream): Automatically intercepts image blocks, routes them to a configured vision model (a catalog provider/model, or a local OpenAI-compatible endpoint), receives a descriptive synthesis, and rewrites the image block into text context[The user attached an image. Description: ...]before handing it to the text-only chat model. - Native Passthrough: Automatically detects models that natively support vision and allows images to pass directly without unnecessary rewriting.
- Rich Visual Tool Suite: Exposes ~40 specialized tools for on-demand OCR (incl. local Tesseract), visual question answering, bounding-box grounding, document/table/formula extraction, QR/barcode reading, UI-flow reconstruction, multi-model consensus and more.
vision_consensusis opt-in via theconsensusEnabledsetting.
🏗️ Architecture
graph LR
User["User attaches Image in Web UI"] --> Gateway["DSH Session Controller"]
Gateway --> BridgeCheck{"Modality Bridge (v0.5.3)"}
BridgeCheck -->|"Augments inputModalities"| SessionAllowed["Prompt Accepted"]
SessionAllowed --> Hook["agent/pre-step Hook"]
Hook --> CheckNative{"Does chat model support vision natively?"}
CheckNative -->|"Yes (Native Passthrough)"| NativeLLM["Send Raw Image to Chat LLM"]
CheckNative -->|"No (Text-Only)"| VisionRouter["Vision Bridge Channels"]
VisionRouter --> VisionModel["Dedicated Vision Model\n(DSH / OpenAI / Ollama / Webhook)"]
VisionModel --> Description["Generated Text Description + OCR"]
Description --> Rewrite["Substitute Image with Text Marker"]
Rewrite --> ChatModel["Send Enriched Text to Chat LLM"]
ChatModel --> Answer["Assistant Response in Chat"]
✨ Key Features
1. Processing Modes
hybrid(default): Automatically describes attached images in chat turns while keeping all ~40 explicit vision tools available for follow-up reasoning.llm: Pure auto-rewrite mode — images are transparently converted to text context; tools remain callable.tools: Auto-rewrite disabled — the chat model is expected to explicitly invokedescribe_imageor OCR tools when required.
2. Multi-Channel Endpoint Routing & Fallback
Chain multiple vision backends with automatic failover, parallel racing, and circuit breaker:
dsh-catalog: Auto-detect or select any vision-capable model already registered in DSH.openai-compatible: Standard OpenAI-compatible vision endpoints (vLLM, SGLang, OpenRouter, etc.).ollama: Auto-discovery and local inference through any vision-capable model the local Ollama instance exposes.webhook/custom: External HTTP or JSON-RPC vision endpoints.
3. High-Performance LRU Description Cache
Caches vision responses by hash(bytes + prompt + model + mode) to eliminate redundant vision API calls and save token quota on repeated questions about the same image.
4. Comprehensive Visual Tool Inventory (~40 Tools)
| Tool Category | Tools | Description |
|---|---|---|
| Core | describe_image, read_image, inspect_image |
General image analysis by attachment ID, file path, or URL. |
| Geometry & Detection | vision_ground, vision_crop, vision_detect, vision_compare, vision_present |
Bounding box coordinates (0–1000 scale), object inventory, multi-image comparison. |
| OCR & Text | vision_ocr, vision_ocr_local, vision_long_ocr, vision_trace, vision_colors, vision_extract_foreground |
Transcription, local Tesseract OCR (offline), long screenshot stitching, SVG tracing, color palettes. |
| Structured & UI | vision_describe_structured, vision_vqa, vision_ui_layout, vision_translate_image |
JSON breakdown ({summary, ocr, layout, entities}), short VQA, UI section analysis. |
| Pixel & Diagnostics | vision_pixel_diff, vision_quality_check |
Semantic visual diff, quality scoring (blur/lighting). |
| Documents & Intelligence | vision_extract_formula, vision_extract_table, vision_scan_barcode, vision_extract_structured, vision_audit_accessibility |
Formula (LaTeX), table (Markdown/HTML), QR/barcode scan, JSON-schema extraction, WCAG accessibility audit. |
| Scenarios, Consensus & Memory | vision_ui_flow, vision_consensus, vision_memory_search |
User-journey graph (Mermaid), multi-model consensus, semantic search over remembered images. |
| Attachments (v0.5.33) | vision_attach_pages, vision_attach_frames, vision_attach_images |
Publish PDF pages, video frames and local/remote images as conversation attachments so native-vision chat models read the pixels themselves. |
📦 Installation
dsh plugin --profile web add @goodandready/dsh-vision-bridge
After installation, restart the DSH Web UI. The configuration card is available under Settings → Plugins → vision-bridge.
⚙️ Configuration (settings.yaml)
dsh-vision-bridge:
# Operation mode: 'hybrid' | 'llm' | 'tools'
mode: hybrid
# Auto-detect vision model or specify provider/model explicitly
visionProvider: ""
visionModel: ""
# Allow vision models to receive images natively without rewriting ('prefer' | 'never' | 'always')
nativePassthrough: prefer
hideRedundantTools: true # hide bridge tools when the chat model already sees images
attachMaxItems: 8 # pages/frames/files one attach call may publish
# Enable LRU description cache
cacheEnabled: true
cacheMaxEntries: 200
# Request timeout in milliseconds
timeoutMs: 120000
# Multi-channel routing configuration
channels: []
channelFallback: sequential # 'sequential' | 'parallel-race'
Parameter Reference
| Parameter | Type | Default | Description |
|---|---|---|---|
mode |
string |
"hybrid" |
Processing mode (hybrid, llm, tools). |
visionProvider |
string |
"" |
ID of the vision provider (empty = auto-detect). |
visionModel |
string |
"" |
ID of the vision model (empty = auto-detect). |
nativePassthrough |
string |
"prefer" |
Behavior for native vision models (prefer, never, always). |
hideRedundantTools |
boolean |
true |
When the chat model supports images natively, hide the compensation tools from that agent and keep only the extra instruments. |
attachMaxItems |
number |
8 |
Maximum images one vision_attach_* call publishes (PDF pages, video frames, files). Hard ceiling 32. |
cacheEnabled |
boolean |
true |
Enables LRU caching for descriptions. |
cacheMaxEntries |
number |
200 |
Maximum number of cached items in memory. |
timeoutMs |
number |
120000 |
Execution timeout in milliseconds. |
channelFallback |
string |
"sequential" |
Channel routing (sequential, parallel-race); ordering via channelOrderMode. |
15. 🛡️ Enterprise Security & Settings Governance (v0.5.27)
- Masked Secrets:
GET /channelsautomatically masks sensitive provider API keys (sk-p...7890or********) and reportshasApiKey: trueto prevent secret leakage in browser DevTools/XHR.POST /channelspreserves existing keys when a mask is submitted. - CSRF & Origin Guard: All mutating and cost-incurring endpoints (
/config,/channels,/upload-pdf,/test,/bench,/batch,DELETE /journal,DELETE /cache, and/doctor?probe=1) validatesec-fetch-site !== 'cross-site'viaisTrustedSettingsRequest, rejecting cross-origin attacks with403 Forbidden. The defaultGET /doctorreport is static (no channel probes). - SSRF fetch policy: model-supplied image URLs are fetched only through
safeFetch— non-http(s) schemes, localhost names and private/loopback/link-local hosts (incl. IPv4-mapped IPv6 and NAT64) are refused on every redirect hop; response bodies are capped. UseallowedUrlHoststo explicitly re-allow an internal endpoint. - apiKeyRef: channels may reference a credential-service entry or environment variable by NAME instead of storing a plaintext
apiKeyin settings; keys resolve at call time.GET /channelskeeps masking values; key preservation on save matches channels by identity, not position. - Privacy & honesty:
maskPII,stripEXIF,auditLogandconsensusEnabledsettings are wired end-to-end; the previously decorative face-blur / NSFW / tiling toggles and no-op presets were removed. - Native settingsScope Integration: Settings UI binds directly to
ctx.settingsScope(namespace: 'dsh-vision-bridge'), supporting reactive snapshot listeners and kernel state synchronization.
📝 Changed in v0.5.30
Security and honesty release. Additive summary of what changed for users:
- SSRF fetch policy: model-supplied image URLs (
describe_imageurls,inspect_image, and the headless-chrome tools) are now fetched only through a policy layer — non-http(s) schemes, localhost names and private/loopback/link-local hosts (incl. IPv4-mapped IPv6 and NAT64) are refused on every redirect hop; bodies are capped. New settingallowedUrlHosts(exact-hostname allowlist) deliberately re-allows an internal endpoint. apiKeyRef: channels can reference a credential-service entry or environment variable BY NAME — plaintext API keys are no longer required insettings.yaml. Existing inlineapiKeyvalues keep working; masked-key preservation on save now matches channels by identity, not position.- Route guards:
POST /bench,POST /batch,DELETE /batch/:id,DELETE /journal,DELETE /cachenow require same-origin like the other mutating routes.GET /doctoris static by default; channel probes run only with?probe=1(same-origin required). - Settings honesty:
maskPII,stripEXIF,auditLogandconsensusEnabledare now wired end-to-end. The previously decorativeblurFaces,nsfwFilter,tileLargeImages/tileThresholdtoggles and the no-op Local/Cloud/LM Studio presets were removed. Changed in v0.5.30: if you relied on them, note they never had an effect. - English source language: all user-facing strings are English; the bundled Russian dictionary was removed — the translation plugin supplies Russian at runtime.
- Core split: the pure kernel (config schema + helpers) moved to
lib/vision-core.js;lib/index.jsre-exports it — no API changes.describe_imagelost a v0.5.13 regression that returned an empty description;vision_annotateworks again; pHash caching no longer mixes up similar images.
📝 Changed in v0.5.31
Maintenance release — no user-facing behavior changes.
- Internal structure: tool registrations moved to
lib/tools/*domain modules (core / grounding / ocr / document / analysis / media);lib/index.jsre-exports everything as before. Smaller files, explicit dependencies between the host and the tool domains. - Settings card hardening: locale registration is fault-tolerant, the redundant sidebar fallback was removed, plugin service access goes through a safe
ctx.getwrapper, and/configaccepts the expanded field set (cacheMaxEntries,channelFallback).
📝 Changed in v0.5.32
Stability and architecture release.
- Internal structure: all ~44 tool registrations moved to
lib/tools/*domain modules (core / grounding / ocr / document / analysis / media) with explicit dependencies;lib/index.jskeeps the host wiring only. - Stability fixes: finished batch records are released after a 10-minute poll window (memory growth fixed);
/upload-pdfrejects payloads above the newmaxPdfBytessetting (20 MiB default) instead of buffering arbitrary bodies;vision_memory_searchscores each attachment against its own description (previously all attachments matched identically); dead host code removed; journal labels are consistent between the legacy and channels paths. - Settings: new
maxPdfBytessetting (upload hard cap).
📝 Changed in v0.5.33
Vision-first release: the bridge now serves chat models that see images natively, not only text-only ones.
- Attach domain (
vision_attach_pages,vision_attach_frames,vision_attach_images): PDF pages, sampled video frames and local/directory/URL images are published as conversation attachments, so a native-vision chat model looks at the pixels itself instead of paying for a second vision call. Every image is compressed by the existingimageMaxWidth/imageMaxHeight/imageQualitysettings and bounded bymaxImageBytes; URL sources go through the same SSRF policy as the rest, and local paths are restricted byallowedImageDirs. - Model-aware tool exposure: on a route whose chat model already accepts images, the compensation tools of the bridge are hidden from that agent (the model does not need them) and only the extra instruments remain; on a text-only route the attach tools are hidden instead, because that model cannot see an attachment. The behaviour is controlled by the new
hideRedundantToolssetting (on by default). - New settings in the plugin card: an Attachments group exposes
attachMaxItems(whole number 1–32, default 8 — how many images one attach call publishes) andhideRedundantTools. Both are validated on save; a value outside the range is rejected with a clear message. The image dimension fields are now also written to the live settings snapshot, not only to the route. - Batch API:
DELETE /batch/:idreleases a finished batch immediately, under the same same-origin guard as start/cancel. The batch record's TTL timer no longer keeps a short-lived process alive, which cut the test suite from 10 minutes to ~3.5 seconds. - Tools mode: images attached in chat are now indexed before the sanitisation gate, so tools that take an attachment id (and the
read_imagealias) work intoolsmode as well; previously the ids were unavailable there. - Fixes: one unreadable source no longer aborts
vision_attach_images— it is reported asSkipped N: <name>: <reason>while the readable sources still attach, and a fetch-policy refusal stays a hard error; the PDF text layer actually appears now (pdftotextwas never detected, so the layer silently never shipped); a page range or frame count cut by the cap is reported throughtruncatedand the note. - Internal: CI installs poppler without
sudo, runs once per commit, serializes per ref and installs ffmpeg for the frame tests; the repository gained the pull-request template and ignores.worktrees/.
📝 Changed in v0.6.0
Major stability, reliability, and lifecycle release.
- In-App Auto-Updater: added dedicated one-click update endpoint (
/api/dsh-vision-bridge/update) and Settings card controls (UpdaterBlock) insettings.plugin.item. Features real-time npm version check, progress state, and restart guidance, protected by loopback and CSRF origin verification gates. - Error Resilience & Catch Elimination: audited and replaced all 76 unannotated empty catch blocks across core runtime and tools. Added
bestEffort(label, fn, fallback)helper for non-fatal side effects, explicit warning propagation in image processing (smartOptimizeImagereturnspreprocessed: booleanandwarnings: string[]), and transparent fallback reporting. - Identity & Scope Integrity: strictly aligned scoped package identity
@goodandready/dsh-vision-bridgeacross manifest, Cordis patch, client module loader, and internal runtime metadata. - Hardened Settings Security: upgraded settings mutation route to a strict fail-closed validator (
isTrustedSettingsRequest) enforcing loopback origin, CSRF header checks, bearer tokens, and rejecting suspicious remote headers. - Native Theme Compliance: client diagnostics and card styles fully migrated to DSH CSS design tokens (
--dsw-alias-*), eliminating all hardcoded color literals. - Sanitized Public Release: integrated plumbing-based release script (
publish.sh) and.gitattributesexport filters ensuring zero private development artifacts in public distribution.
📝 Changed in v0.6.2
Reliability hardening, error visibility, and degradation warnings release.
- Real bestEffort Logging: replaced all remaining dummy
/* bestEffort ... */ void errcatches across core runtime, tools, and routes with activebestEffort(label, fn)logging toconsole.debug. - Degradation Warnings in Tools: all tool modules now declare
warnings: { type: 'array', items: { type: 'string' } }in output schemas and surfacewarnings: string[]on partial JSON parsing errors or engine fallbacks. - Image Preprocessing Failure Capture: preprocessing exceptions (deskew, enhance, stripEXIF, compression) are recorded and surfaced to tool callers instead of failing silently.
- Expanded Verification: added unit test suite
test/issue-314-besteffort-warnings.test.js, bringing total test coverage to 329 passing tests across 104 suites.
📄 License
MIT © GooDAnDReaDY
Links
More in this category
liustack/modlens★ 4183
Vision bridge for text-only models: paste an image, get structured JSON evidence (OCR, layout, semantics).
ysr666/dsh-vision-router★ 1137
Free vision for text-only agents: built-in keyless vision chain plus pixel tools (Q&A, grounding, crop, pixel diff, colors, OCR, SVG trace, cutout, screenshots); paste an image to use it.
Anionex/dsh-vision-toolkit★ 886
Vision for text-only models: paste an image and the model switches to a Vision Toolkit variant for image Q&A, multi-image comparison, long-screenshot OCR, screenshot-to-UI reproduction, element grounding, and pixel diff. No API key by default — images are processed by the author-hosted free service, 100 per machine per day; configurable to your own provider.
fandc520/dsh-comfyui★ 105
Drive a local or remote ComfyUI server from DeepSeek Harness: comfyui_run / comfyui_object_info / comfyui_workflow tools generate and edit images and videos, with a workflow library (graph extraction: per component / main flow / all), a load area with resolution auto-match, a live queue, SDXL and Wan 2.1 templates, a companion skill, and a same-origin media proxy.
dickpy/dsh-imagegen★ 101
AI image generation for the DSH Web GUI: text-to-image and image-to-image through a configurable OpenAI-compatible endpoint (gpt-image-2 / gpt-image-1 / dall-e-3), with an api_url/api_key settings card and a sidebar split-pane generation studio.
sunxin-ai/dsh-design-qa★ 44
Design-fidelity QA for text-only models: a `deepseek_vision` tool borrows an eye from any OpenAI-compatible vision route, so the model can judge whether an implementation matches its mock — shipped with the benchmark behind that judgement (four fixtures, 23 injected defects, raw transcripts) and the questioning discipline it depends on.
Community comments
Comments are public GitHub Discussions. Loading them connects to GitHub and Giscus; a GitHub account is required to post.