Plug-in vision for text-only models on DSH, with native interaction for image understanding and generation, GUI automation, through layered evidence memory and cache.
Install
# from npm (prebuilt)
dsh plugin --profile web add dsh-mindseye
# from GitHub (first run asks for allowBuilds approval — follow the hint, retry)
dsh plugin --profile web add github:kanchengw/dsh-mindseye
Any plugin you install runs third-party code with your own permissions — it can read your files, use your credentials, and reach the network, and tool approvals don’t sandbox it. GitHub-sourced plugins also run build scripts at install time — pnpm blocks those until you allow them, so an install can stop with ERR_PNPM_GIT_DEP_PREPARE_NOT_ALLOWED or ERR_PNPM_IGNORED_BUILDS; dsh prints the exact key to add under allowBuilds in your profile’s pnpm-workspace.yaml, and the install works on the next run. Allowing a build is a trust decision: only install sources you trust, and pin a commit (github:owner/repo#sha).
README

Intent-driven vision, image generation, and visible browser automation for DeepSeek Harness.
MindsEye is a plugin for DeepSeek Harness. It gives text-only models access to image understanding, image generation, and optional browser automation while keeping the DSH conversation as the main user experience.
Capabilities
Image understanding
- Preserves DSH image attachments instead of asking users to select local files manually.
- Automatically mounts vision tools on image turns. Text-only turns keep a single activation entry until vision is needed.
mindseye_read_imagehandles visual questions and focused tasks such as OCR, layout, charts, colors, and pixel differences.mindseye_groundreturns a target's pixel bounding box for downstream actions such as clicking or cropping.- Supports single-image and multi-image reads, with structured results for images, evidence, answers, and call metadata.
Image generation and editing
mindseye_generate_imagesends a user's image request to the configured image-generation route.mindseye_edit_imagesends a DSH image attachment and an edit request to the configured image-editing route.- Generated images are returned as native DSH attachments and displayed in the conversation.
- Generation does not automatically save files to the project or run a verification pass.
Browser automation
When gui.enabled is turned on, MindsEye opens a separate visible Chrome or Edge session. The GUI tools can open pages, take snapshots, wait, click, type, send key presses, scroll, and close the session.
If a page requires CAPTCHA, login, or permission confirmation, the run pauses on a native DSH question card. The user can:
- take over the visible browser and complete the step;
- skip the first handoff question when the step may already be complete; or
- abandon the run.
After the user resumes, MindsEye checks the page state before returning control to the model. The browser uses an isolated session and does not attach to the user's existing Chrome or Edge profile. GUI actions require a fresh snapshot after each action so element references and coordinates cannot silently become stale.
Tools
| Tool | Purpose |
|---|---|
mindseye_plan |
Extracts the current request and prepares the intent context used by downstream tools. |
mindseye_read_image |
Answers questions about one or more images and extracts focused visual evidence. |
mindseye_ground |
Locates a target and returns its pixel bounding box. |
mindseye_generate_image |
Generates an image from the user's request. |
mindseye_edit_image |
Edits a supplied image attachment. |
mindseye_vision_activate |
Mounts the vision tools during a text-only turn. |
mindseye_gui_open / snapshot / wait |
Opens a browser session and observes its current state. |
mindseye_gui_click / type / keypress / scroll |
Performs a state-checked browser action. |
mindseye_gui_close |
Closes the current browser session. |
The memory tools are optional and expose explicit DSH operations for storing, retrieving, searching, and comparing image-related records.
Configuration
Configure MindsEye from the DSH settings card or the plugin configuration.
vision.routes: independent routes forunderstand,extract, andlocate.vision.fallbacks: fallback routes for vision calls.image.generate: ordered image-generation routes.image.edit: ordered image-editing routes.gui.enabled: enables the visible browser tools. It is disabled by default.gui.browser:auto,chrome, oredge.gui.restrictHosts: enables host allowlisting when set totrue.gui.allowedHosts: hosts allowed when host restriction is enabled.gui.maxStepsandgui.timeoutMs: limits for one browser run.
Vision routes use OpenAI-compatible Chat Completions or Responses APIs. Image routes support JSON and multipart request bodies so different image providers can be configured independently.
Data and Safety
- Image-capable models keep native DSH image blocks. Text-only fallback paths use isolated temporary files created for the current paste operation.
- Image bytes and questions are sent to the configured provider only when a MindsEye tool makes that provider call.
- Credentials come from DSH credentials, environment variables, or plugin settings and are sent only to the matching provider.
- Browser automation starts a local Chrome or Edge child process only when explicitly enabled. It does not use the user's existing browser profile and does not execute downloaded code.
- Browser navigation can be restricted to an explicit host allowlist.
Install
npm install dsh-mindseye
npx @deepseek-ai/dsh plugin --profile web add dsh-mindseye
Restart DSH Web after installation. Then configure at least one vision route in the MindsEye settings card. Unconfigured focused vision routes fall back to the general understanding route when available.
Development
pnpm install
pnpm test
pnpm typecheck
pnpm build
Links
More in this category
liustack/modlens★ 4100
Vision bridge for text-only models: paste an image, get structured JSON evidence (OCR, layout, semantics).
ysr666/dsh-vision-router★ 1128
Free vision for text-only agents: built-in keyless vision chain plus pixel tools (Q&A, grounding, crop, pixel diff, colors, OCR, SVG trace, cutout, screenshots); paste an image to use it.
Anionex/dsh-vision-toolkit★ 887
Vision for text-only models: paste an image and the model switches to a Vision Toolkit variant for image Q&A, multi-image comparison, long-screenshot OCR, screenshot-to-UI reproduction, element grounding, and pixel diff. No API key by default — images are processed by the author-hosted free service, 100 per machine per day; configurable to your own provider.
dickpy/dsh-imagegen★ 99
AI image generation for the DSH Web GUI: text-to-image and image-to-image through a configurable OpenAI-compatible endpoint (gpt-image-2 / gpt-image-1 / dall-e-3), with an api_url/api_key settings card and a sidebar split-pane generation studio.
fandc520/dsh-comfyui★ 94
Drive a local or remote ComfyUI server from DeepSeek Harness: comfyui_run / comfyui_object_info / comfyui_workflow tools generate and edit images and videos, with a workflow library (graph extraction: per component / main flow / all), a load area with resolution auto-match, a live queue, SDXL and Wan 2.1 templates, a companion skill, and a same-origin media proxy.
sunxin-ai/dsh-design-qa★ 44
Design-fidelity QA for text-only models: a `deepseek_vision` tool borrows an eye from any OpenAI-compatible vision route, so the model can judge whether an implementation matches its mock — shipped with the benchmark behind that judgement (four fixtures, 23 injected defects, raw transcripts) and the questioning discipline it depends on.
Community comments
Comments are public GitHub Discussions. Loading them connects to GitHub and Giscus; a GitHub account is required to post.