Agent-callable vision tool that describes local images via any OpenAI-compatible vision endpoint you configure, with an optional multi-model cross-check and no built-in keys.
Install
# from GitHub (first run asks for allowBuilds approval — follow the hint, retry)
dsh plugin --profile web add github:TZHR-invest/dsh-plugins#path:/packages/dsh-vision
Any plugin you install runs third-party code with your own permissions — it can read your files, use your credentials, and reach the network, and tool approvals don’t sandbox it. GitHub-sourced plugins also run build scripts at install time — pnpm blocks those until you allow them, so an install can stop with ERR_PNPM_GIT_DEP_PREPARE_NOT_ALLOWED or ERR_PNPM_IGNORED_BUILDS; dsh prints the exact key to add under allowBuilds in your profile’s pnpm-workspace.yaml, and the install works on the next run. Allowing a build is a trust decision: only install sources you trust, and pin a commit (github:owner/repo#sha).
README
A vision tool for DeepSeek Harness agents: sends a local image to the OpenAI-compatible vision endpoint you configure and returns a text description — for when the current session model can't ingest images directly.
The plugin ships no built-in defaults: provider (baseURL), credentials, and models are all configured at install time. If configuration is incomplete the plugin loads normally but does not register the tool (log line shows what's missing) — it can never crash the host.
Features
- Agent-callable tool:
vision image_path=/tmp/shot.png - Any OpenAI-compatible
/chat/completionsendpoint (vendor-agnostic) - Credential resolution chain: inline
apiKey→apiKeyEnv(environment variable) →~/.dsh/.credentials.yamlsame-name key - Optional cross-check: query several models and merge answers (
cross_check=true) to guard against hallucinated descriptions maxTokens/maxImageByteslimits (optional)- Requests go out via a
curlsubprocess — inherits host proxy env vars and avoids Cloudflare 403/1010 on default urllib/undici user agents
Install
Tarball
tar xzf dsh-vision-tool-install.tar.gz && cd dsh-vision-tool-install
# interactive: just run it and answer the prompts
bash install.sh --restart
# or parameterized (scriptable / repeatable deploys)
bash install.sh --restart \
--vision-base-url https://example.com/v1 \
--vision-api-key-env MY_VISION_KEY \
--vision-model gpt-4o \
--vision-models gpt-4o,qwen-vl-max,kimi-latest \
--vision-max-tokens 2000
npm
dsh plugin --profile web add dsh-vision-tool
# then configure: edit the dsh-vision-tool config section in ~/.dsh/profiles/web/cordis.patch.yml
Install parameters (--vision-*)
| Parameter | Required | Meaning |
|---|---|---|
--vision-base-url |
✅ | OpenAI-compatible chat/completions endpoint |
--vision-model |
✅ | Default vision model |
--vision-api-key |
one of | API key inline |
--vision-api-key-env |
one of | Environment variable name (also tried as a key in ~/.dsh/.credentials.yaml) |
--vision-models |
optional | Cross-check model list (comma-separated; needed for cross_check=true) |
--vision-max-tokens |
optional | Omitted from the request if not set |
No parameters + interactive terminal → guided prompts. No parameters + non-interactive → empty config written (tool not registered); re-run with parameters, or edit the config section in
~/.dsh/profiles/web/cordis.patch.yml.
Usage (inside the agent)
vision image_path=/tmp/shot.png
vision image_path=/tmp/shot.png question="What text is in this image?"
vision image_path=/tmp/shot.png model=gpt-4o
vision image_path=/tmp/shot.png cross_check=true # requires --vision-models configured at install
Security
- No built-in keys — credentials come only from your install parameters / environment / credential file
~/.dsh/.credentials.yamlshould be chmod 600
Troubleshooting
- Tool doesn't appear: log shows
[dsh-vision-tool] missing ...— re-run install.sh with--vision-*parameters and restart - Credential resolution fails: at least one of
apiKey/apiKeyEnv(incl. credential-file same-name key) - Endpoint incompatible: baseURL must be OpenAI-compatible (
/chat/completions,Authorization: Bearer) - Empty output from reasoning models: configure
--vision-max-tokens 2000or higher
License
MIT. Chinese documentation: README.zh.md
Part of dsh-plugins — a small monorepo of DSH plugins: dsh-lan-gateway, dsh-vision-tool, dsh-mobile-ui.
Links
More in this category
liustack/modlens★ 4063
Vision bridge for text-only models: paste an image, get structured JSON evidence (OCR, layout, semantics).
ysr666/dsh-vision-router★ 1125
Free vision for text-only agents: built-in keyless vision chain plus pixel tools (Q&A, grounding, crop, pixel diff, colors, OCR, SVG trace, cutout, screenshots); paste an image to use it.
Anionex/dsh-vision-toolkit★ 884
Vision for text-only models: paste an image and the model switches to a Vision Toolkit variant for image Q&A, multi-image comparison, long-screenshot OCR, screenshot-to-UI reproduction, element grounding, and pixel diff. No API key by default — images are processed by the author-hosted free service, 100 per machine per day; configurable to your own provider.
dickpy/dsh-imagegen★ 93
AI image generation for the DSH Web GUI: text-to-image and image-to-image through a configurable OpenAI-compatible endpoint (gpt-image-2 / gpt-image-1 / dall-e-3), with an api_url/api_key settings card and a sidebar split-pane generation studio.
fandc520/dsh-comfyui★ 89
Drive a local or remote ComfyUI server from DeepSeek Harness: comfyui_run / comfyui_object_info / comfyui_workflow tools generate and edit images and videos, with a workflow library (graph extraction: per component / main flow / all), a load area with resolution auto-match, a live queue, SDXL and Wan 2.1 templates, a companion skill, and a same-origin media proxy.
sunxin-ai/dsh-design-qa★ 44
Design-fidelity QA for text-only models: a `deepseek_vision` tool borrows an eye from any OpenAI-compatible vision route, so the model can judge whether an implementation matches its mock — shipped with the benchmark behind that judgement (four fixtures, 23 injected defects, raw transcripts) and the questioning discipline it depends on.
Community comments
Comments are public GitHub Discussions. Loading them connects to GitHub and Giscus; a GitHub account is required to post.