Renders pasted/uploaded images inline in the DSH Web chat and gives text-only models vision: the model-invokable visual_review tool calls any OpenAI-compatible multimodal API first, falling back to a local Qwen3-VL worker.
Install
# from npm (prebuilt)
dsh plugin --profile web add visual-review
# from GitHub (first run asks for allowBuilds approval — follow the hint, retry)
dsh plugin --profile web add github:wang-bool/visual-review
Any plugin you install runs third-party code with your own permissions — it can read your files, use your credentials, and reach the network, and tool approvals don’t sandbox it. GitHub-sourced plugins also run build scripts at install time — pnpm blocks those until you allow them, so an install can stop with ERR_PNPM_GIT_DEP_PREPARE_NOT_ALLOWED or ERR_PNPM_IGNORED_BUILDS; dsh prints the exact key to add under allowBuilds in your profile’s pnpm-workspace.yaml, and the install works on the next run. Allowing a build is a trust decision: only install sources you trust, and pin a commit (github:owner/repo#sha).
README
visual-review
A dual-sided plugin for the DeepSeek Harness (DSH) Web UI: render images in the chat interface and let the model interpret them.
- Image rendering — Images pasted or uploaded by the user (PNG / JPEG / WebP / GIF) show up directly inside the chat bubble.
- Visual interpretation — The
visual_reviewtool calls a vision-language model and returns a Chinese description of the image (text, objects, people, scenes, charts, …). - Dual engines — Cloud-first: any OpenAI-compatible multimodal
chat/completionsAPI, zero local dependencies; falls back to a local Qwen3-VL-8B worker when no cloud config is set (data never leaves the machine). - No model swap needed — The plugin rewrites
imageblocks into text annotations carrying the attachment ID on the send path, so any text-only model can "see" images by calling the tool.
Table of Contents
- Features
- Architecture & How It Works
- Directory Layout
- Requirements
- Installation
- Configuration
- Usage
- Security Notes
- Development & Testing
- FAQ
- License
Features
| Capability | Description |
|---|---|
| Render images in chat | The client plugin injects a conversation.chat.node render slot and renders image blocks as <img> (bytes served by the /vr-image route) |
| Let the model "see" images | /api/session.prompt interception + agent/pre-step fallback: image blocks → annotation text with the attachment ID, instructing the model to call the tool |
| Image interpretation tool | visual_review accepts attachment_id (session image) or image_path (local file) |
| Cloud engine | OpenAI-compatible chat/completions; the image is inlined as a data: URL; handles content / reasoning_content replies |
| Local engine | Built-in JSON-lines worker running a local Qwen3-VL-8B; images and results stay on the machine |
| Config tool | visual_review_config reads / updates the cloud url / api_key / model and switches engines in one step |
Architecture & How It Works
┌──────────────────────────── DSH Web frontend ──────────────────────────┐
│ Chat UI (React) │
│ └─ visual-review client plugin (lib/client.js) │
│ injects conversation.chat.node render slots │
│ └─ image block → <img src="/vr-image/sha256:xxx"> │
└───────────────┬─────────────────────────────────────────┬──────────────┘
│ ① /api/session.prompt (image→annotation)│ ② /vr-image bytes
┌───────────────▼─────────────────────────────────────────▼──────────────┐
│ DSH Host (plugin host side lib/index.js) │
│ • /api/session.prompt intercept route: save image as attachment, │
│ replace it with annotation text │
│ • agent/pre-step fallback (idempotent): image blocks in model-side │
│ messages → annotations │
│ • /vr-image route: loopback-only, reads image bytes from attachment │
│ storage │
│ • visual_review tool: cloud first → local fallback │
│ ├─ cloud: inline node script → OpenAI-compatible chat/completions│
│ └─ local: spawn Python worker (server/visual_review_server.py) │
└────────────────────────────────────────────────────────────────────────┘
Three key mechanisms:
Send interception (image → annotation). When the client issues
session.prompt, the plugin intercepts it, saves eachimageblock through the attachments service to obtain an attachment ID (e.g.sha256:xxx), and replaces the block with an annotation like:[用户上传了图片附件(image/png,640x360px)。你无法直接查看图片, 请调用 visual_review 工具并传入 attachment_id="sha256:xxx" 获取图片的视觉解析。]The
agent/pre-stephook applies the same transformation to model-side messages as an idempotent fallback. This is how a text-only model can "see" images — it just follows the annotation and calls the tool./vr-imageroute (bytes → rendering). The client render slot extracts the attachment ID from the annotation and renders<img src="/vr-image/<id>">; the host reads the bytes back from attachment storage. The route only accepts loopback (127.0.0.1 / ::1) requests.visual_reviewtool (image → description). With the attachment ID in hand, the model calls the tool. It reads$DSH_HOME/visual_review.config.jsonfor cloud settings; otherwise it spawns/reuses the local Python worker and sends the image as base64 over stdin.
Local worker protocol (JSON-lines, one JSON object per line):
request → {"id":1, "kind":"describe", "image_b64":"...", "prompt":"...", "max_new_tokens":256}
reply → {"id":1, "ok":true, "text":"图片中有一只橘猫……"}
ready → {"event":"ready"} (model loads in a background thread; describe waits for it)
Directory Layout
visual-review/
├── lib/
│ ├── index.js # Host plugin: tools, routes, interception, worker mgmt
│ └── client.js # Client plugin: image rendering in the chat UI
├── server/
│ └── visual_review_server.py # Local inference worker (Qwen3-VL-8B, JSON-lines)
├── config/
│ └── visual_review.config.example.json # Cloud config template (API-key placeholder)
├── install/
│ ├── install.sh # One-shot installer into a DSH profile
│ └── cordis.patch.yml.example # Snippet to append to cordis.patch.yml manually
├── examples/
│ └── test_image.png # Sample image for testing
├── assets/
│ ├── qr-official-account.jpg # Official-account QR code
│ └── qr-group.png # Community group QR code
├── cordis.patch.yml # dsh.bundle install patch (used by `dsh plugin add`)
├── package.json # Plugin manifest (host + client entries, dsh.bundle)
├── README.md / README.en.md # 中文 / English docs
├── LICENSE # MIT
└── .gitignore
Requirements
- DSH (DeepSeek Harness) Web profile (
~/.dsh/profiles/web), Node.js ≥ 18 (globalfetchrequired). - Vision engine, pick one:
- Cloud mode (recommended, simplest): any OpenAI-compatible multimodal
chat/completionsAPI (SiliconFlow, OpenAI, DeepSeek, local vLLM, …). - Local mode: Python 3.10+,
torch,transformers(≥ 4.57 for Qwen3-VL),Pillow, and local weights ofQwen3-VL-8B-Instruct(any compatible Qwen3-VL checkpoint works).
- Cloud mode (recommended, simplest): any OpenAI-compatible multimodal
Installation
Options A/B need no external distribution channel. The plugin declares a
dsh.bundlemanifest (cordis.patch.yml), so the officialdsh plugin addflow works too (Option C, recommended).
Option C: official dsh plugin add (recommended)
After the npm release:
dsh plugin --profile web add visual-review
Not published yet, or want the latest from source? Install straight from GitHub:
dsh plugin --profile web add github:wang-bool/visual-review
Then restart dsh web and refresh the browser.
Option A: one-shot script
git clone <this repo> visual-review
cd visual-review
bash install/install.sh
The script copies the plugin into $DSH_HOME/profiles/web/node_modules/visual-review/ and registers it in cordis.patch.yml (idempotent — safe to re-run). Override DSH_HOME / DSH_PROFILE for other data dirs or profiles.
Option B: manual install
Put the package into the profile's
node_modules:DSH_HOME="${DSH_HOME:-$HOME/.dsh}" PROFILE_DIR="$DSH_HOME/profiles/web" mkdir -p "$PROFILE_DIR/node_modules/visual-review" cp -r package.json lib server "$PROFILE_DIR/node_modules/visual-review/"Append to
$PROFILE_DIR/cordis.patch.yml(seeinstall/cordis.patch.yml.example):- insert: - id: visual-review name: 'visual-review'Restart the DSH web service and refresh the browser. The host log should show:
visual_review: 全局插件已加载 (config=.../visual_review.config.json) visual_review: /vr-image 图片路由已注册 visual_review: /api/session.prompt 拦截路由已注册
Configuration
Cloud engine (recommended)
Write $DSH_HOME/visual_review.config.json (default ~/.dsh/visual_review.config.json):
{
"url": "https://api.siliconflow.cn/v1/chat/completions",
"apiKey": "sk-xxxx",
"model": "Qwen/Qwen3.5-35B-A3B"
}
Alternatively, ask the model to call the visual_review_config tool with url / api_key / model (only the fields you pass are changed; passing none clears the cloud config and falls back to the local engine).
This file holds the API key in plain text — keep it out of git (the repo's
.gitignorealready covers it) and mind file permissions.
Local engine
Falls back automatically when no cloud config exists. You need local weights:
huggingface-cli download Qwen/Qwen3-VL-8B-Instruct
Model dir resolution: CLI argument → env var VISUAL_REVIEW_MODEL_DIR → default ~/.cache/huggingface/hub/models--Qwen--Qwen3-VL-8B-Instruct.
Environment variables
| Variable | Default | Description |
|---|---|---|
DSH_HOME |
$HOME/.dsh |
DSH data dir; where visual_review.config.json lives |
VISUAL_REVIEW_SERVER |
plugin's own server/visual_review_server.py |
path to the local inference script |
VISUAL_REVIEW_PYTHON |
python3 |
interpreter for the local worker |
VISUAL_REVIEW_WORKSPACE |
sandbox workspace root → process cwd | workspace root (where .visual_review_tmp lives) |
VISUAL_REVIEW_ATTACH_ROOT |
$HOME/.dsh/attachments/v1/objects |
attachment object-store root (used by /vr-image fallback) |
VISUAL_REVIEW_MODEL_DIR |
default HF cache path | local Qwen3-VL model dir (read by the worker) |
Usage
- Try it directly: paste or upload an image in the chat → it renders immediately in the bubble; the annotation appears on the model side and the model calls
visual_reviewon its own. - Analyze a local file: ask the model, e.g. “please analyze
/path/to/photo.pngwithvisual_review” — the model passes it viaimage_path. - Custom questions: be specific, e.g. “read all the text in this image” or “describe the trend of this chart”. The tool accepts a
promptparameter the model forwards as needed. - Switch engines: update the cloud config via
visual_review_config; clearing all three fields returns to the local engine.
Security Notes
- The
/vr-imageand/api/session.promptroutes only accept loopback (127.0.0.1 / ::1); other sources get 403. - The cloud API key is stored in plain text at
$DSH_HOME/visual_review.config.json— never commit it; restrict file permissions. - Local mode: images are processed on the machine only; no external network requests are made.
Development & Testing
Test the local worker standalone (without DSH)
# via env var
VISUAL_REVIEW_MODEL_DIR=/path/to/Qwen3-VL-8B-Instruct python3 -u server/visual_review_server.py
# via CLI argument
python3 -u server/visual_review_server.py /path/to/Qwen3-VL-8B-Instruct
From another terminal, send a request (base64-encode the image first):
B64=$(base64 -w0 examples/test_image.png)
printf '{"id":1,"kind":"describe","image_b64":"%s","prompt":"请描述这张图片","max_new_tokens":256}\n' "$B64" | \
VISUAL_REVIEW_MODEL_DIR=/path/to/model python3 -u server/visual_review_server.py
Expect {"event":"ready"} followed by {"id":1,"ok":true,"text":"..."}. kind:"decode_image" decodes a base64 file back to binary without loading the model.
Protocol cheatsheet
| Direction | Message |
|---|---|
| Host → worker | {"id":<n>,"kind":"describe","image_b64":"...","prompt":"...","max_new_tokens":<int>} |
| Host → worker | {"id":<n>,"kind":"decode_image","b64_path":"...","out_path":"..."} |
| worker → Host | {"id":<n>,"ok":true,"text":"..."} or {"id":<n>,"ok":false,"error":"..."} |
| worker → Host | {"event":"ready"} / {"event":"load-failed","error":"..."} |
FAQ
Q: The model loads slowly / “模型加载超时(20 分钟)”.
Local mode must load the Qwen3-VL-8B weights on first use — can take minutes depending on disk/GPU. Check VISUAL_REVIEW_MODEL_DIR or switch to cloud mode.
Q: “找不到附件 ID”.
Attachment references are only valid within the current session: re-upload the image, or pass a local path via image_path.
Q: Cloud call fails, hinting the url seems to miss /chat/completions.
url must be a full OpenAI-compatible endpoint, e.g. https://xxx/v1/chat/completions.
Q: The image shows “(图片加载失败)” in the chat.
/vr-image only allows loopback requests; remote access to the DSH Web UI may be refused. Also, during development the client plugin requires rebuilding the frontend bundle and a refresh.
Q: After pasting an image the model says it can't see it.
Check the host log for /api/session.prompt 拦截路由已注册. If missing (webServer/apiProxy unavailable), interception is skipped and only the agent/pre-step fallback applies; if it still fails, re-upload the image.
License
MIT © visual-review authors
Contact
Welcome to follow the official account or join the community group for latest news, feedback, and open-source discussion:
| Official account · Latest news | Community group · Chat & feedback |
|---|---|
💡 The group QR code expires after 7 days; follow the official account to get the latest invite link.
Links
More in this category
liustack/modlens★ 4100
Vision bridge for text-only models: paste an image, get structured JSON evidence (OCR, layout, semantics).
ysr666/dsh-vision-router★ 1128
Free vision for text-only agents: built-in keyless vision chain plus pixel tools (Q&A, grounding, crop, pixel diff, colors, OCR, SVG trace, cutout, screenshots); paste an image to use it.
Anionex/dsh-vision-toolkit★ 887
Vision for text-only models: paste an image and the model switches to a Vision Toolkit variant for image Q&A, multi-image comparison, long-screenshot OCR, screenshot-to-UI reproduction, element grounding, and pixel diff. No API key by default — images are processed by the author-hosted free service, 100 per machine per day; configurable to your own provider.
dickpy/dsh-imagegen★ 99
AI image generation for the DSH Web GUI: text-to-image and image-to-image through a configurable OpenAI-compatible endpoint (gpt-image-2 / gpt-image-1 / dall-e-3), with an api_url/api_key settings card and a sidebar split-pane generation studio.
fandc520/dsh-comfyui★ 94
Drive a local or remote ComfyUI server from DeepSeek Harness: comfyui_run / comfyui_object_info / comfyui_workflow tools generate and edit images and videos, with a workflow library (graph extraction: per component / main flow / all), a load area with resolution auto-match, a live queue, SDXL and Wan 2.1 templates, a companion skill, and a same-origin media proxy.
sunxin-ai/dsh-design-qa★ 44
Design-fidelity QA for text-only models: a `deepseek_vision` tool borrows an eye from any OpenAI-compatible vision route, so the model can judge whether an implementation matches its mock — shipped with the benchmark behind that judgement (four fixtures, 23 injected defects, raw transcripts) and the questioning discipline it depends on.
Community comments
Comments are public GitHub Discussions. Loading them connects to GitHub and Giscus; a GitHub account is required to post.