Reuses DeepSeek web's built-in vision mode for text-only models: the deepseek_vision tool drives the local deepseek-vision-cli browser automation (manual login helper, deep-think enabled, auto-closes browser) and returns image descriptions as text.
Install
# from GitHub (first run asks for allowBuilds approval — follow the hint, retry)
dsh plugin --profile web add github:Cheng-cheng9669/dsh-deepseek-vision
Any plugin you install runs third-party code with your own permissions — it can read your files, use your credentials, and reach the network, and tool approvals don’t sandbox it. GitHub-sourced plugins also run build scripts at install time — pnpm blocks those until you allow them, so an install can stop with ERR_PNPM_GIT_DEP_PREPARE_NOT_ALLOWED or ERR_PNPM_IGNORED_BUILDS; dsh prints the exact key to add under allowBuilds in your profile’s pnpm-workspace.yaml, and the install works on the next run. Allowing a build is a trust decision: only install sources you trust, and pin a commit (github:owner/repo#sha).
README
A DSH plugin that gives text-only DeepSeek models image recognition.
It registers the deepseek_vision tool, which calls the local
deepseek-vision-cli
(browser automation driving chat.deepseek.com's vision mode) and returns the
image description as text to the model.
Tool
deepseek_vision(image, prompt?)image: absolute path to a local image (PNG/JPG/WebP/GIF; HEIC is not supported)prompt: optional question or analysis request for the image
Installation
dsh plugin --profile web add D:/Dsh/tools/dsh-deepseek-vision
Or use the release tarball:
dsh plugin --profile web add dsh-deepseek-vision-0.0.1.tgz
First login
DeepSeek web vision requires a web login session. The current web UI uses password / third-party login, so the recommended way is the manual login helper:
py -3.13 D:\Dsh\tools\deepseek-vision-cli\dsv_manual_login.py
It opens an Edge window to the DeepSeek login page. After you log in manually,
the helper saves the token to D:\Dsh\.cache\dsv_token.
Token login (alternative)
If you already have a DeepSeek web session in your normal browser, you can copy the token directly:
- Open
https://chat.deepseek.comand log in. - Press
F12→Application→Local Storage→https://chat.deepseek.com. - Find
userToken, copy itsvaluefield. - Save it to
D:\Dsh\.cache\dsv_token(no quotes, no newline):
[IO.File]::WriteAllText(
'D:\Dsh\.cache\dsv_token',
'PASTE_TOKEN_HERE',
(New-Object Text.UTF8Encoding($false))
)
The token is only stored locally in D:\Dsh\.cache\dsv_token; it is never
committed to this repository.
Configuration
Default configuration for this machine:
pythonCommand:pypythonVersion:-3.13dsvScript:D:/Dsh/tools/deepseek-vision-cli/dsv.pytimeoutMs:180000maxOutputChars:20000
These can be overridden in a profile patch.
Security notes
- The tool only accepts local file paths and invokes the external CLI through
the DSH
subprocessservice with an argv array, never through a shell. - Images are sent to DeepSeek's web endpoint; this is not local recognition.
- This relies on non-official web interfaces and may be subject to risk control. Use a secondary account and keep usage low-frequency.
- Static security audit reports
criticalfor spawning a subprocess; this is inherent to wrapping an external CLI. The code has been manually reviewed: it only runs the fixeddsv.py, with no command concatenation or extra exfiltration.
License
MIT
Links
More in this category
liustack/modlens★ 4100
Vision bridge for text-only models: paste an image, get structured JSON evidence (OCR, layout, semantics).
ysr666/dsh-vision-router★ 1128
Free vision for text-only agents: built-in keyless vision chain plus pixel tools (Q&A, grounding, crop, pixel diff, colors, OCR, SVG trace, cutout, screenshots); paste an image to use it.
Anionex/dsh-vision-toolkit★ 887
Vision for text-only models: paste an image and the model switches to a Vision Toolkit variant for image Q&A, multi-image comparison, long-screenshot OCR, screenshot-to-UI reproduction, element grounding, and pixel diff. No API key by default — images are processed by the author-hosted free service, 100 per machine per day; configurable to your own provider.
dickpy/dsh-imagegen★ 99
AI image generation for the DSH Web GUI: text-to-image and image-to-image through a configurable OpenAI-compatible endpoint (gpt-image-2 / gpt-image-1 / dall-e-3), with an api_url/api_key settings card and a sidebar split-pane generation studio.
fandc520/dsh-comfyui★ 94
Drive a local or remote ComfyUI server from DeepSeek Harness: comfyui_run / comfyui_object_info / comfyui_workflow tools generate and edit images and videos, with a workflow library (graph extraction: per component / main flow / all), a load area with resolution auto-match, a live queue, SDXL and Wan 2.1 templates, a companion skill, and a same-origin media proxy.
sunxin-ai/dsh-design-qa★ 44
Design-fidelity QA for text-only models: a `deepseek_vision` tool borrows an eye from any OpenAI-compatible vision route, so the model can judge whether an implementation matches its mock — shipped with the benchmark behind that judgement (four fixtures, 23 injected defects, raw transcripts) and the questioning discipline it depends on.
Community comments
Comments are public GitHub Discussions. Loading them connects to GitHub and Giscus; a GitHub account is required to post.