Screen capture and external vision recognition: take_screenshot, list_windows, analyze_image and view_image tools with a configurable GPT vision channel (gpt-5.5 / gpt-5.6-sol / gpt-5.6-terra), API key via the credentials service, and a settings card; view_image shows the screenshot in the Web UI while the model context keeps text only.
Install
# from GitHub (first run asks for allowBuilds approval — follow the hint, retry)
dsh plugin --profile web add github:ankye/dsh-client-vision#path:/packages/tool-vision
Any plugin you install runs third-party code with your own permissions — it can read your files, use your credentials, and reach the network, and tool approvals don’t sandbox it. GitHub-sourced plugins also run build scripts at install time — pnpm blocks those until you allow them, so an install can stop with ERR_PNPM_GIT_DEP_PREPARE_NOT_ALLOWED or ERR_PNPM_IGNORED_BUILDS; dsh prints the exact key to add under allowBuilds in your profile’s pnpm-workspace.yaml, and the install works on the next run. Allowing a build is a trust decision: only install sources you trust, and pin a commit (github:owner/repo#sha).
README
English | 中文
Give your DeepSeek Harness agent eyes. dsh-client-vision is a screen-capture + external image-recognition plugin for DeepSeek Harness: the agent takes a screenshot (or points at any image), hands it to a vision-capable model through a pluggable channel, and gets back plain text it can actually act on — no multimodal model required.
Compatibility
This revision requires DeepSeek Harness core 0.1.2-alpha.5 or later within the 0.1.x line. It uses the Settings service API introduced in that core release.
Why you want it
- DeepSeek can't see — now it can. The harness model has no image input. This plugin runs the whole "look" outside the model and returns text the agent can reason about, exactly like Codex's semantic vision tool.
- Capture anything, any way.
fullscreen/window(with live window enumeration) /region/interactive— grab the browser, a game window, or one corner of the screen. - Multi-channel by design. Tools are decoupled from recognition backends. The
gptchannel ships ready to use; adding Claude, Gemini, or a local model is oneanalyze()implementation + one registry line — the three tools never change. - Secret-safe. The API key lives in the harness
credentialsstore (VISION_GPT_API_KEY) — never in settings files, logs, or the conversation transcript. - Every preset, out of the box. Mounted on the host plane, so
code,standard,cordis,minimal— every agent sees the tools. No preset switching. - Ready to ship. Prebuilt bundles included; three install paths (drop into the monorepo /
pnpm publish/ tarball). - Smart payloads. Large captures are auto-downscaled and re-encoded (≤1568px JPEG q80) before they leave the machine.
Capabilities
Tools
| Tool | What it does |
|---|---|
take_screenshot |
Capture the screen: fullscreen (primary display), window (by id from list_windows), region (x, y, width, height), interactive (user selection), android (adb device/emulator), or ios (booted simulator). Returns the PNG path + dimensions. |
list_windows |
Enumerate on-screen windows (id, app, title) — macOS CGWindowList, Windows Get-Process main handles, Linux X11 (wmctrl/xprop) — pick the browser or game window to capture. |
analyze_image |
Submit an image (a path, or the most recent screenshot) to the configured vision channel and return a plain-text description. |
view_image |
One-shot "look at this": capture the screen (or use image_path) and recognize it through the active channel. The screenshot is rendered as an image card in the Web conversation, while the model context receives only the plain-text description — the image bytes never enter the model context. |
Platforms
| Platform | Capture backend | Window enumeration | Extra requirements |
|---|---|---|---|
| macOS | screencapture (system) |
Swift CGWindowList |
Screen Recording permission on first use |
| Windows | PowerShell System.Drawing (system) |
Get-Process main window handles |
PowerShell System.Drawing |
| Linux | ImageMagick import |
wmctrl + xprop |
ImageMagick (convert/identify), wmctrl, x11-utils |
mode=interactive (system selection UI) is macOS-only; on Windows and Linux
use mode=region with explicit coordinates.
Device capture
| Mode | What it captures | Requirements |
|---|---|---|
android |
A connected Android device or emulator screen | adb on PATH with a device online (adb devices); works from any host. With several devices online, pass device=<serial>. |
ios |
The booted iOS simulator | macOS host with Xcode (xcrun simctl) |
Settings (vision namespace)
Configured in Settings → Plugins → Plugin configuration → Vision:
| Field | Meaning |
|---|---|
Endpoint (baseUrl) |
Domain + optional path prefix; /chat/completions is appended. e.g. https://api.example.com/v1 |
| Channel | The active recognition backend (currently gpt). |
| Model | gpt-5.5 / gpt-5.6-sol / gpt-5.6-terra |
| API key | Stored through the harness credentials service as VISION_GPT_API_KEY; the literal never leaves your machine. |
Channels
| Channel | Backend | Model | API key |
|---|---|---|---|
gpt |
OpenAI-compatible /chat/completions |
gpt-5.5 / gpt-5.6-sol / gpt-5.6-terra |
required (e.g. VISION_GPT_API_KEY) |
zhipu |
Zhipu GLM-4V, OpenAI-compatible /chat/completions |
glm-4v-plus / glm-4v-flash |
required (e.g. VISION_ZHIPU_API_KEY) |
ollama |
local Ollama /api/chat (default http://localhost:11434) |
llava / llava-llama3 / bakllava / moondream / qwen2-vl / minicpm-v (or any installed vision model) |
none |
Pick the channel in Settings → Plugins → Vision; the model dropdown follows
the channel and the API-key control is hidden for ollama. For ollama the
base URL defaults to http://localhost:11434 and the model to llava when
left blank.
Multi-channel architecture
model → analyze_image(image, prompt)
│ reads vision.channel
▼
channels/<id>/analyze() ← one implementation per backend
│
gpt: POST {baseUrl}/chat/completions (image_url data URL)
claude / gemini / local: … ← add yours here
Adding a channel is deliberately small:
// src/channels/<id>/index.ts
export async function myAnalyze(ctx, call): Promise<string> {
// call.imageB64, call.mime, call.prompt, call.config, call.signal
return await fetchYourVisionApi(...)
}
// src/channels/index.ts — one registry line
export const channels = {
gpt: { label: 'GPT', analyze: gptAnalyze },
myChannel: { label: 'My Channel', analyze: myAnalyze },
}
The tools (take_screenshot / list_windows / analyze_image) and their schemas never change.
Installation (official — no repo modification)
dsh plugin add installs the packages into your profile; each package declares dsh.bundle, so the rows mount automatically — no patch rows, no repo edits.
Prerequisites
- DeepSeek Harness core
0.1.2-alpha.5or later within the0.1.xline, plusdshandpnpmon PATH.
1. Get the packages (pick one)
a. From this repository (recommended until published to npm):
dsh plugin --profile web add \
file:/path/to/dsh-client-vision/packages/tool-vision \
file:/path/to/dsh-client-vision/packages/ui-vision
b. Tarball:
cd packages/tool-vision && npm pack
cd packages/ui-vision && npm pack
dsh plugin --profile web add file:/path/to/deepseek-ai-dsh-tool-vision-0.1.0-rc.7.tgz \
file:/path/to/deepseek-ai-dsh-client-ui-vision-0.1.0-rc.7.tgz
c. npm registry (after publishing):
dsh plugin --profile web add @deepseek-ai/dsh-tool-vision @deepseek-ai/dsh-client-ui-vision
A
[WARN] Issues with peer dependenciesmessage is expected and safe to ignore — the peers come from your deployment's own bundles at runtime.
2. Verify
node -e "console.log(JSON.stringify(require(process.env.HOME + '/.dsh/profiles/web/package.json').dsh.profile.bundles))"
# should list dsh-tool-vision and dsh-client-ui-vision
3. Restart + configure
Restart the harness, then Settings → Plugins → Plugin configuration → Vision: set the endpoint, model, and your own API key (VISION_GPT_API_KEY), save.
4. Verify
Ask the agent to "look at the screen" — it should call take_screenshot → analyze_image and describe what it sees.
Uninstall
dsh plugin --profile web remove @deepseek-ai/dsh-tool-vision @deepseek-ai/dsh-client-ui-vision
Alternative: build inside a harness fork
If you run a fork of deepseek-harness (not the official deployment), you can drop the packages into the monorepo instead:
cp -R packages/tool-vision <harness>/packages/vision/tool-vision
cp -R packages/ui-vision <harness>/packages/client/ui-vision
Then add both to apps/cli/package.json (workspace:^), add ./packages/vision/tool-vision to tsconfig.host.json and ./packages/client/ui-vision to tsconfig.client.json, pnpm install, build (tsdown host + client passes), and restart.
Quick start
- Restart the harness.
- The tool catalog now includes
take_screenshot/list_windows/analyze_image. - Open Settings → Plugins → Plugin configuration → Vision, set the endpoint, model, and your own API key, and save.
- Ask the agent to "look at the screen" — it will screenshot and describe what it sees.
Development
- This repository is a source distribution: the peer packages (
@deepseek-ai/dsh-tools, …) resolve from your deployment.lib/ships prebuilt, sonpm packworks immediately. - The
tsconfig.jsonfiles are standalone; the harness monorepo's build pipeline (including the client-bundletsdown.config.ts) applies in Option A. - Never commit secrets. The API key stays in each machine's
.credentials.yaml.
License
MIT
Links
More in this category
liustack/modlens★ 4121
Vision bridge for text-only models: paste an image, get structured JSON evidence (OCR, layout, semantics).
ysr666/dsh-vision-router★ 1126
Free vision for text-only agents: built-in keyless vision chain plus pixel tools (Q&A, grounding, crop, pixel diff, colors, OCR, SVG trace, cutout, screenshots); paste an image to use it.
Anionex/dsh-vision-toolkit★ 884
Vision for text-only models: paste an image and the model switches to a Vision Toolkit variant for image Q&A, multi-image comparison, long-screenshot OCR, screenshot-to-UI reproduction, element grounding, and pixel diff. No API key by default — images are processed by the author-hosted free service, 100 per machine per day; configurable to your own provider.
dickpy/dsh-imagegen★ 99
AI image generation for the DSH Web GUI: text-to-image and image-to-image through a configurable OpenAI-compatible endpoint (gpt-image-2 / gpt-image-1 / dall-e-3), with an api_url/api_key settings card and a sidebar split-pane generation studio.
fandc520/dsh-comfyui★ 97
Drive a local or remote ComfyUI server from DeepSeek Harness: comfyui_run / comfyui_object_info / comfyui_workflow tools generate and edit images and videos, with a workflow library (graph extraction: per component / main flow / all), a load area with resolution auto-match, a live queue, SDXL and Wan 2.1 templates, a companion skill, and a same-origin media proxy.
sunxin-ai/dsh-design-qa★ 44
Design-fidelity QA for text-only models: a `deepseek_vision` tool borrows an eye from any OpenAI-compatible vision route, so the model can judge whether an implementation matches its mock — shipped with the benchmark behind that judgement (four fixtures, 23 injected defects, raw transcripts) and the questioning discipline it depends on.
Community comments
Comments are public GitHub Discussions. Loading them connects to GitHub and Giscus; a GitHub account is required to post.