Gives text-only models vision: forwards user images to an OpenAI-compatible vision model and shows the descriptions in a Web UI right panel.
Install
# from GitHub (first run asks for allowBuilds approval — follow the hint, retry)
dsh plugin --profile web add github:jyh20030112/dsh-visual-plugin
Any plugin you install runs third-party code with your own permissions — it can read your files, use your credentials, and reach the network, and tool approvals don’t sandbox it. GitHub-sourced plugins also run build scripts at install time — pnpm blocks those until you allow them, so an install can stop with ERR_PNPM_GIT_DEP_PREPARE_NOT_ALLOWED or ERR_PNPM_IGNORED_BUILDS; dsh prints the exact key to add under allowBuilds in your profile’s pnpm-workspace.yaml, and the install works on the next run. Allowing a build is a trust decision: only install sources you trust, and pin a commit (github:owner/repo#sha).
README
A plugin for DeepSeek Harness.
Features
- Native image understanding — uploaded images stay on DSH's native attachment and model path; the plugin does not configure or call a separate vision model.
- Copyable image history — the right panel records the current DSH model's final answer beside each image thumbnail, with expandable history and one-click copy.
- Plugin-owned video upload — accepts MP4, M4V, MOV, AVI, MPG/MPEG, MKV, and WebM only when extension, signature, and FFprobe agree.
- Scene-aware video analysis — normalizes to H.264/yuv420p MP4, extracts keyframes with PySceneDetect, and sends ordered timestamped images to the current DSH vision model.
- Right-side panel — switch between image/video views, play normalized videos directly, and stage a selected video in the chat draft.
- Advanced video settings — tune upload size, storage quota, duration, output size, FPS, CRF, and keyframe count from the plugin settings card.
How it works
image → DSH native attachment → current image-capable model → final answer
→ /vision-bridge/recent → panel thumbnail + copyable description
video → container validation → H.264/yuv420p normalization → PySceneDetect
→ timestamped keyframes → DSH native image attachments → current model answers
The plugin never rewrites model messages or calls a private vision endpoint. Select an image-capable model in DSH before sending images or asking about a video.
Quick start
Video support requires FFmpeg/FFprobe >= 6.1 from the same major release (with libx264) and PySceneDetect >= 0.7.1 < 0.8 installed on the host:
ffmpeg -version
ffprobe -version
python -m pip install 'scenedetect[opencv]>=0.7.1,<0.8'
scenedetect version
The plugin never downloads these tools or runs installers. Image features remain available when they are missing, and the settings card reports each video dependency issue.
dsh plugin --profile web add dsh-visual-plugin # or: github:jyh20030112/dsh-visual-plugin
When developing this checkout against a local DeepSeek Harness source tree, install the local package instead:
cd /absolute/path/to/dsh-visual-plugin
npm run bootstrap
dsh plugin --profile web add link:/absolute/path/to/dsh-visual-plugin
bootstrap automatically finds a sibling or ancestor-adjacent Harness checkout.
For another layout, set its location explicitly:
HARNESS=/absolute/path/to/deepseek-harness npm run bootstrap
Restart dsh web, then:
Open Settings → Plugins → Plugin configuration and expand the Visual Media card. Use Sidebar to show or hide the right panel, and adjust the advanced video settings when needed.
Select an image-capable model in DSH; there is no separate vision-model configuration in this plugin.
Send an image. The current model answers natively, and the image panel records the thumbnail and final answer for copying.
Upload a video from Upload video beside the composer. Once processing finishes, select Videos in the right panel to play it; Ask in chat stages a draft and never submits automatically.
Vision model
Image and keyframe understanding use the image-capable model currently selected in DSH. Model providers, endpoints, and credentials are managed by DSH rather than this plugin.
Uninstall
dsh plugin --profile web remove dsh-visual-plugin
Restart dsh web. The command forwards to pnpm remove inside the profile, and the bundle layer list reconciles to drop the plugin automatically.
Project layout
src/
index.ts native image history, video_describe tool, settings, and HTTP routes
config.ts advanced video-processing settings and runtime policy
video/ upload, container probing, transcoding, scene detection, keyframes, and HTTP Range playback
client/ image/video panel, upload controls, advanced settings, locales, and CSS
cordis.patch.yml bundle patch layer
Build
npm run bootstrap && npm run typecheck && npm run build # needs a local harness checkout
Prebuilt lib/ is committed, so consumers never build.
CI/CD
ci.yml verifies artifacts and the pack contents on every push/PR. release.yml (tag v*) checks the version, packs, creates a GitHub Release, and publishes to npm.
Resources
- DeepSeek Harness — the plugin host this project extends.
- PySceneDetect — scene detection used to select video keyframes.
- awesome-dsh-plugin — the curated DSH plugin list where this plugin is registered.
Friendly Links
Thanks
- HsiangNianian — for their help and insights during development.
- tingfeng347 — for the build-stability and local-harness-setup fixes.
- dsh-auto-continue — a DSH Web UI plugin that auto-resumes interrupted requests with 「继续」 (error classification, adaptive backoff, browser notifications); a handy companion.
License
Links
More in this category
zhu1090093659/dsh-web-ui#packages/dsh-tool-describe-image★ 8121
A `describe_image` vision tool for text-only models: images (local path, URL, attachment) go to a configurable OpenAI-compatible vision endpoint and only the returned text enters the session.
liustack/modlens★ 4063
Vision bridge for text-only models: paste an image, get structured JSON evidence (OCR, layout, semantics).
ysr666/dsh-vision-router★ 1125
Free vision for text-only agents: built-in keyless vision chain plus pixel tools (Q&A, grounding, crop, pixel diff, colors, OCR, SVG trace, cutout, screenshots); paste an image to use it.
Anionex/dsh-vision-toolkit★ 884
Vision for text-only models: paste an image and the model switches to a Vision Toolkit variant for image Q&A, multi-image comparison, long-screenshot OCR, screenshot-to-UI reproduction, element grounding, and pixel diff. No API key by default — images are processed by the author-hosted free service, 100 per machine per day; configurable to your own provider.
dickpy/dsh-imagegen★ 93
AI image generation for the DSH Web GUI: text-to-image and image-to-image through a configurable OpenAI-compatible endpoint (gpt-image-2 / gpt-image-1 / dall-e-3), with an api_url/api_key settings card and a sidebar split-pane generation studio.
fandc520/dsh-comfyui★ 89
Drive a local or remote ComfyUI server from DeepSeek Harness: comfyui_run / comfyui_object_info / comfyui_workflow tools generate and edit images and videos, with a workflow library (graph extraction: per component / main flow / all), a load area with resolution auto-match, a live queue, SDXL and Wan 2.1 templates, a companion skill, and a same-origin media proxy.
Community comments
Comments are public GitHub Discussions. Loading them connects to GitHub and Giscus; a GitHub account is required to post.