Free vision bridge for text-only models: image understanding, OCR, UI and debug analysis via free-tier providers (Qwen3-VL-Flash, Doubao, DeepSeek-OCR) with a settings GUI.
Install
# from npm (prebuilt)
dsh plugin --profile web add dsh-free-vision
# from GitHub (first run asks for allowBuilds approval — follow the hint, retry)
dsh plugin --profile web add github:FuzzySoul/dsh-free-vision
Any plugin you install runs third-party code with your own permissions — it can read your files, use your credentials, and reach the network, and tool approvals don’t sandbox it. GitHub-sourced plugins also run build scripts at install time — pnpm blocks those until you allow them, so an install can stop with ERR_PNPM_GIT_DEP_PREPARE_NOT_ALLOWED or ERR_PNPM_IGNORED_BUILDS; dsh prints the exact key to add under allowBuilds in your profile’s pnpm-workspace.yaml, and the install works on the next run. Allowing a build is a trust decision: only install sources you trust, and pin a commit (github:owner/repo#sha).
README
This plugin publishes its README in Chinese only.
🌐 English | 中文
DSH 免费视觉插件 — 让纯文本模型获得看图能力(截图、报错、UI 分析、OCR、文档),优先使用各平台免费视觉模型,零 MCP 配置。
Free vision plugin for DeepSeek Harness (dsh) — image understanding for text-only models using free-tier vision models, with zero MCP configuration.
为什么免费 / Why free
默认使用免费额度充足的提供商,无账单惊吓:
| 提供商 | 模型 | 免费额度 | API Key 环境变量 |
|---|---|---|---|
| qwen(默认) | Qwen3-VL-Flash | 阿里云百炼限免(激活送 50万 token) | DASHSCOPE_API_KEY |
| volcengine | 豆包视觉模型 | 火山引擎豆包免费 token(20万起,可申请 50万) | VOLCENGINE_API_KEY |
| siliconflow | DeepSeek-OCR | 硅基流动 OCR 免费 | SILICONFLOW_API_KEY |
| zhipu | GLM-4.6V | 按量 | ZHIPU_API_KEY |
| hunyuan | HY-Vision | 按量 | HUNYUAN_API_KEY |
| custom | 任意 OpenAI 兼容 | — | CUSTOM_API_KEY + CUSTOM_BASE_URL + CUSTOM_MODEL_NAME |
一张 1MB 截图 ≈ 2600 token,qwen 限免额度可分析约 19 万张图。 One 1MB screenshot ≈ 2,600 tokens ≈ $0.0006 on qwen; free quota covers ~190,000 images.
特性 / Features
- 零 MCP 配置 — 不用改
cordis.patch.yml、运行时不用npx:视觉引擎(luma-mcp)作为本包依赖内置,进程内启动 - 单个通用工具 —
image_understand(可用config.toolName改名)注册到ctx.tools,每次请求模型都能看到 - 免费优先、多提供商 — 千问 / 豆包 / 硅基流动免费档开箱即用;智谱 / 混元 / custom 可切换
- 每个 Provider 可覆盖 API Base URL — 内置 Provider 可指向代理、API Gateway、本地服务或任意 OpenAI 兼容端点,无需改成 custom
- 直连 — 子进程剥离代理环境变量,国内 API 直连(带代理会导致 502)
- 任务模式 —
auto | general | ocr | ui | debug | describe;大图自动多裁剪保真 - 中英双语 — 工具描述与文档中英文都可用
安装 / Install
dsh plugin --profile web add dsh-free-vision
重启 dsh web 后,工具 image_understand 即可用。
Restart dsh web; the tool appears as image_understand.
设置界面 / Settings UI
重启 dsh web 后,打开 设置 → Free Vision 即可看到配置表单(API Key、提供商、
工具名等),由插件 schema 自动渲染。保存到 ~/.dsh/free-vision.json,下一次
调用立即生效,无需重启。
After restart, open Settings → Free Vision — a form for every config option,
saved to ~/.dsh/free-vision.json, effective on the next tool call.
配置 / Configuration
- id: free-vision
name: 'dsh-free-vision'
config:
apiKey: 'sk-xxxx' # 可选:缺省回退到提供商环境变量
baseURLs: {} # 可选:按 Provider 覆盖 API Base URL,例如 { qwen: 'https://my-proxy.example.com/v1' }
modelProvider: qwen # qwen | volcengine | siliconflow | zhipu | hunyuan | custom
modelName: qwen3-vl-flash # 可选模型覆盖
toolName: image_understand # 工具公开名(冲突时可改名)
maxTokens: 8192
temperature: 0.7
multiCrop: true
toolCallTimeoutMs: 200000
lumaEnv: {} # 传递给视觉引擎的额外环境变量
也可以只设置对应的环境变量(如 DASHSCOPE_API_KEY)。
Or just set the matching environment variable (e.g. DASHSCOPE_API_KEY).
覆盖 API 地址 / Base URL override
当 baseURLs 缺失或值为空时,继续使用该 Provider 的官方默认地址。
| Provider | 默认 Base URL |
|---|---|
| qwen | https://dashscope.aliyuncs.com/compatible-mode/v1 |
| volcengine | https://ark.cn-beijing.volces.com/api/v3 |
| siliconflow | https://api.siliconflow.cn/v1 |
| zhipu | https://open.bigmodel.cn/api/paas/v4 |
| hunyuan | https://api.hunyuan.cloud.tencent.com/v1 |
引擎会自动拼接 /chat/completions 并避免重复路径,因此以下写法都可用:
https://my-proxy.example.com/v1https://my-proxy.example.com/v1/chat/completions
也可以使用环境变量:QWEN_BASE_URL、VOLCENGINE_BASE_URL、
SILICONFLOW_BASE_URL、ZHIPU_BASE_URL、HUNYUAN_BASE_URL
(以及 custom 的 CUSTOM_BASE_URL)。
免费 Key 申请 / Free API keys
| 提供商 | 免费 Key 获取 |
|---|---|
| qwen | 阿里云百炼 bailian.console.aliyun.com — 开通即送免费额度,模型选 qwen3-vl-flash(限免) |
| volcengine | 火山引擎 volcengine.com — 豆包新用户送免费 token(20万起,可申请 50万) |
| siliconflow | 硅基流动 siliconflow.cn — DeepSeek-OCR 免费调用 |
用法 / Usage
模型调用 image_understand 时传入:
image_source(必填):本地路径、HTTP(S) URL 或 data URI(PNG/JPG/WebP/GIF,≤10MB)prompt(必填):对图片的问题 — 中英文均可task_type(可选):auto | general | ocr | ui | debug | describe
工作原理 / How it works
dsh web → cordis 加载 free-vision → 进程内启动视觉引擎(版本锁定)
→ MCP 连接 → 注册 image_understand 到 ctx.tools
→ 模型调用工具 → 引擎预处理(压缩 / 多裁剪)→ 免费视觉 API(直连)
→ 返回文字证据
开发 / Development
npm install
node test-plugin.mjs # 端到端冒烟测试(需要 API Key 环境变量)
许可证 / License
MIT — 封装 luma-mcp(MIT)与 MCP SDK(MIT)。免费额度数据来自各平台官方页面,可能变动,使用前请核实。
Links
More in this category
liustack/modlens★ 4063
Vision bridge for text-only models: paste an image, get structured JSON evidence (OCR, layout, semantics).
ysr666/dsh-vision-router★ 1125
Free vision for text-only agents: built-in keyless vision chain plus pixel tools (Q&A, grounding, crop, pixel diff, colors, OCR, SVG trace, cutout, screenshots); paste an image to use it.
Anionex/dsh-vision-toolkit★ 884
Vision for text-only models: paste an image and the model switches to a Vision Toolkit variant for image Q&A, multi-image comparison, long-screenshot OCR, screenshot-to-UI reproduction, element grounding, and pixel diff. No API key by default — images are processed by the author-hosted free service, 100 per machine per day; configurable to your own provider.
dickpy/dsh-imagegen★ 93
AI image generation for the DSH Web GUI: text-to-image and image-to-image through a configurable OpenAI-compatible endpoint (gpt-image-2 / gpt-image-1 / dall-e-3), with an api_url/api_key settings card and a sidebar split-pane generation studio.
fandc520/dsh-comfyui★ 89
Drive a local or remote ComfyUI server from DeepSeek Harness: comfyui_run / comfyui_object_info / comfyui_workflow tools generate and edit images and videos, with a workflow library (graph extraction: per component / main flow / all), a load area with resolution auto-match, a live queue, SDXL and Wan 2.1 templates, a companion skill, and a same-origin media proxy.
sunxin-ai/dsh-design-qa★ 44
Design-fidelity QA for text-only models: a `deepseek_vision` tool borrows an eye from any OpenAI-compatible vision route, so the model can judge whether an implementation matches its mock — shipped with the benchmark behind that judgement (four fixtures, 23 injected defects, raw transcripts) and the questioning discipline it depends on.
Community comments
Comments are public GitHub Discussions. Loading them connects to GitHub and Giscus; a GitHub account is required to post.