Model-facing `vision` tool for DeepSeek Harness: describe and OCR image files by calling the free Zhipu GLM vision API directly (glm-4v-flash fallback chain), no external CLI required.
Install
# from npm (prebuilt)
dsh plugin --profile web add dsh-vision-free-eyes
# from GitHub (first run asks for allowBuilds approval — follow the hint, retry)
dsh plugin --profile web add github:314857493/dsh-vision#path:/packages/vision-tool
Any plugin you install runs third-party code with your own permissions — it can read your files, use your credentials, and reach the network, and tool approvals don’t sandbox it. GitHub-sourced plugins also run build scripts at install time — pnpm blocks those until you allow them, so an install can stop with ERR_PNPM_GIT_DEP_PREPARE_NOT_ALLOWED or ERR_PNPM_IGNORED_BUILDS; dsh prints the exact key to add under allowBuilds in your profile’s pnpm-workspace.yaml, and the install works on the next run. Allowing a build is a trust decision: only install sources you trust, and pin a commit (github:owner/repo#sha).
README
This plugin publishes its README in Chinese only.
给 DeepSeek Harness(DSH)的纯文本模型补上免费「眼睛」:一个分析已知本地图片路径的
vision(image, question) 工具,
直连智谱 GLM 免费视觉 API(glm-4v-flash → glm-4.6v-flash → glm-4.1v-thinking-flash 自动降级链),
不依赖任何额外安装的 CLI。完整说明见仓库根目录的
README。
安装(发布后)
dsh plugin --profile web add dsh-vision-free-eyes
或在 profile 的 cordis.patch.yml 中追加:
- insert:
- id: dsh-vision-free-eyes
name: dsh-vision-free-eyes
前提
- DSH
0.1.0-rc.8、0.1.1-rc.1或0.1.1-rc.2。 - 必须:智谱 GLM 免费 key,环境变量
GLM_API_KEY或ZHIPU_API_KEY(open.bigmodel.cn 注册即得,格式id.secret;Windows 也可setx,插件自动读注册表)。 - 出站 HTTPS 到
open.bigmodel.cn。
插件通过 DSH 注入的 tools 服务注册标准 ToolDefinition,不安装或直接导入
@deepseek-ai/dsh-tools 等官方运行时包。
权限与证据边界
- 只读取调用者明确给出的一个绝对图片路径,并只向固定的智谱 HTTPS 端点发送图片字节。
- 只读取 GLM 环境变量;Windows 环境变量缺失时使用固定参数执行系统自带的只读
reg query, 不使用 shell 字符串,也不记录或持久化 Key。 - 无 npm 运行依赖和安装期生命周期脚本;只新增
dsh-vision-free-eyesEntry ID,不写 DSH Profile 或替换官方组件。完整边界见 SECURITY.md。 - 已验证三个声明版本的一次性 Web Profile 安装、配置合成、冷启动和卸载;rollback、真实用户 Profile、带真实 GLM Key 的端到端结果、Windows 运行和独立安全审计仍未验证。下一验收门槛是 推送固定 Commit 后取得公开 CI 运行记录;这些低层证据不会被表述成真实 Profile 或安全审计通过。
使用
告诉模型单个图片文件的已知绝对路径即可,例如:"看一下 D:\xxx\screenshot.png"。默认 image 模式使用
完整 GLM 视觉语言模型理解图片并回答 question;只有用户明确要求逐字提取时才用 mode="ocr"。
no_cache=true 可跳过进程内结果缓存。
工具会在联网前强制检查绝对路径、单文件类型和图片文件魔数;目录、相对路径以及非 png/jpeg/webp/gif/bmp 内容会直接拒绝。模型不应在调用前用 shell 预检路径;目录错误是停止条件, 即使目录里似乎只有一张图片也不会自行遍历,不会把普通本地文件伪装成图片上传。 图片理解会保持用户问题的范围和详细程度,只陈述能够确认的可见事实,并区分总数、当前项与 额外/折叠项等计数语义;简要概述不会逐项转写正文,界面分析也会区分地址栏、搜索框等区域。 单一事实问题只返回所问事实和必要限定,不主动附加未经询问的元素位置或上下文。
该工具不会自动解析 GUI 粘贴/上传附件,也不应遍历 DSH 附件目录猜测图片路径。GUI 贴图请使用 「… + 自动识图」路由;路由已经提供图片描述时,模型应直接使用描述,不要再次调用本工具。 输出为图片内容的纯文字描述,不附加内部耗时标记。
配置
| config | 默认 | 说明 |
|---|---|---|
apiKeyEnv |
GLM_API_KEY / ZHIPU_API_KEY |
GLM key 的环境变量名(可传数组) |
no_cache |
false |
跳过进程内结果缓存,强制重新请求 |
限制
- 单图 ≤ 15MB;支持 png/jpg/jpeg/webp/gif/bmp。
- 不支持 PDF(
doc模式已移除);图片字节会上传到智谱服务器。
Links
More in this category
liustack/modlens★ 4096
Vision bridge for text-only models: paste an image, get structured JSON evidence (OCR, layout, semantics).
ysr666/dsh-vision-router★ 1129
Free vision for text-only agents: built-in keyless vision chain plus pixel tools (Q&A, grounding, crop, pixel diff, colors, OCR, SVG trace, cutout, screenshots); paste an image to use it.
Anionex/dsh-vision-toolkit★ 887
Vision for text-only models: paste an image and the model switches to a Vision Toolkit variant for image Q&A, multi-image comparison, long-screenshot OCR, screenshot-to-UI reproduction, element grounding, and pixel diff. No API key by default — images are processed by the author-hosted free service, 100 per machine per day; configurable to your own provider.
dickpy/dsh-imagegen★ 98
AI image generation for the DSH Web GUI: text-to-image and image-to-image through a configurable OpenAI-compatible endpoint (gpt-image-2 / gpt-image-1 / dall-e-3), with an api_url/api_key settings card and a sidebar split-pane generation studio.
fandc520/dsh-comfyui★ 90
Drive a local or remote ComfyUI server from DeepSeek Harness: comfyui_run / comfyui_object_info / comfyui_workflow tools generate and edit images and videos, with a workflow library (graph extraction: per component / main flow / all), a load area with resolution auto-match, a live queue, SDXL and Wan 2.1 templates, a companion skill, and a same-origin media proxy.
sunxin-ai/dsh-design-qa★ 44
Design-fidelity QA for text-only models: a `deepseek_vision` tool borrows an eye from any OpenAI-compatible vision route, so the model can judge whether an implementation matches its mock — shipped with the benchmark behind that judgement (four fixtures, 23 injected defects, raw transcripts) and the questioning discipline it depends on.
Community comments
Comments are public GitHub Discussions. Loading them connects to GitHub and Giscus; a GitHub account is required to post.