Adds a configurable vision model to text-only main models: a vision_read_image tool, a composer-bar vision-model selector, and automatic image-to-text conversion for text-only routes.
Install
# from GitHub (first run asks for allowBuilds approval — follow the hint, retry)
dsh plugin --profile web add github:poiuyjie/dsh-vision-opencode
Any plugin you install runs third-party code with your own permissions — it can read your files, use your credentials, and reach the network, and tool approvals don’t sandbox it. GitHub-sourced plugins also run build scripts at install time — pnpm blocks those until you allow them, so an install can stop with ERR_PNPM_GIT_DEP_PREPARE_NOT_ALLOWED or ERR_PNPM_IGNORED_BUILDS; dsh prints the exact key to add under allowBuilds in your profile’s pnpm-workspace.yaml, and the install works on the next run. Allowing a build is a trust decision: only install sources you trust, and pin a commit (github:owner/repo#sha).
README
dsh-vision-opencode
DeepSeek can't "see" images? OpenCode multimodal workaround: give your text-only main model a configurable vision model.
What it does
- Image pasted in chat → the vision model (e.g. MiMo-V2.5) converts it to text first, then your main model replies as usual — no model switch needed
- "Vision model" dropdown next to the input box, auto-listing every image-capable model across your providers
- Manage models under Settings → Vision;
vision_read_imagetool /vision-image-analysisskill for OCR, charts, screenshots - Resilience: 60s per-call timeout, 1 retry, and a placeholder fallback so the turn never dies
Install
Option 1 (native DSH, recommended):
dsh plugin --profile web add -w github:poiuyjie/dsh-vision-opencode
Option 2 — one-click script (install.sh Ubuntu / install.ps1 Windows):
curl -fsSL https://raw.githubusercontent.com/poiuyjie/dsh-vision-opencode/main/scripts/install.sh | bash
Restart dsh, then pick a vision model in the dropdown.
Uninstall: dsh plugin --profile web remove -w dsh-vision-opencode (or uninstall.sh).
Back up image conversations before uninstalling — old image chats may no longer reach a text-only main model afterwards.
Configuration
Edit ~/.dsh/settings.yaml (or use Settings → Vision):
vision-opencode:
provider: '' # vision model provider; empty = not chosen
model: '' # vision model id; empty = not chosen
autoConvert: true # auto-convert toggle; set false if problems
Text-only main routes are auto-detected and get image conversion; native multimodal routes keep the stock DSH path.
Turning off reasoning for the vision model
Each model has a "Reasoning" row in Settings → Vision:
| Option | Meaning |
|---|---|
| Default | Follow the provider's default level — think normally |
| Off | No thinking — faster and cheaper. Only when the provider really declares off (e.g. hy3 off:"none") |
| Force off | Best-effort at disabling thinking (e.g. reasoning_effort:"none"), not guaranteed; models without a declared off (like MiMo) get this |
Very few models have a real "Off" entry; the rest show "Default / Force off" with a "not guaranteed" note. Turning thinking off usually cuts first-token latency and cost (MiMo is verified to stop thinking with reasoning_effort:"none").
⚠️ Providers declare "turning off thinking" very inconsistently (
off:"none"/off:null/ no field). The plugin can only tell "Off" from "Force off" per each catalog and best-effort a disabling param — no guarantee every provider can actually disable thinking.
FAQ
- Disable auto-convert only (keep tool + selector): set
vision-opencode.autoConvert: falseand restart - Conversion fails / selector missing: usually no vision model picked or a version mismatch — check the browser console and file an issue
- Text-only vs native multimodal routes are auto-distinguished; no config sync on provider/model switches
Development
- Always tag before pushing: create and push a version tag for every push (e.g.
git tag v0.4.0 && git push origin v0.4.0) so every remote update carries a traceable version marker.
License
MIT
Links
More in this category
liustack/modlens★ 4100
Vision bridge for text-only models: paste an image, get structured JSON evidence (OCR, layout, semantics).
ysr666/dsh-vision-router★ 1128
Free vision for text-only agents: built-in keyless vision chain plus pixel tools (Q&A, grounding, crop, pixel diff, colors, OCR, SVG trace, cutout, screenshots); paste an image to use it.
Anionex/dsh-vision-toolkit★ 887
Vision for text-only models: paste an image and the model switches to a Vision Toolkit variant for image Q&A, multi-image comparison, long-screenshot OCR, screenshot-to-UI reproduction, element grounding, and pixel diff. No API key by default — images are processed by the author-hosted free service, 100 per machine per day; configurable to your own provider.
dickpy/dsh-imagegen★ 99
AI image generation for the DSH Web GUI: text-to-image and image-to-image through a configurable OpenAI-compatible endpoint (gpt-image-2 / gpt-image-1 / dall-e-3), with an api_url/api_key settings card and a sidebar split-pane generation studio.
fandc520/dsh-comfyui★ 94
Drive a local or remote ComfyUI server from DeepSeek Harness: comfyui_run / comfyui_object_info / comfyui_workflow tools generate and edit images and videos, with a workflow library (graph extraction: per component / main flow / all), a load area with resolution auto-match, a live queue, SDXL and Wan 2.1 templates, a companion skill, and a same-origin media proxy.
sunxin-ai/dsh-design-qa★ 44
Design-fidelity QA for text-only models: a `deepseek_vision` tool borrows an eye from any OpenAI-compatible vision route, so the model can judge whether an implementation matches its mock — shipped with the benchmark behind that judgement (four fixtures, 23 injected defects, raw transcripts) and the questioning discipline it depends on.
Community comments
Comments are public GitHub Discussions. Loading them connects to GitHub and Giscus; a GitHub account is required to post.