基于 macOS Vision Framework 的全本地识图插件:`ocr_image`(文字提取,表格结构+坐标)与 `view_image`(场景/人脸/二维码);web 输入框可直接粘贴多张图片,图片永不离开你的 Mac。
安装
# GitHub 源码(首次需按提示配置 allowBuilds 构建授权后重试)
dsh plugin --profile web add github:niyongsheng/free-vision-skill
装任何插件都等于在你的机器上跑第三方代码,权限和你本人一样大——能读你的文件、用你的凭据、访问网络,工具审批管不到它。GitHub 来源的插件还会在安装时执行构建脚本。请只安装可信来源,并尽量锁定 commit(github:owner/repo#sha)。
README
Fully-local image understanding (OCR / table extraction / description) via macOS Vision. Images never leave your Mac.
Install (DSH-Plugin)
dsh plugin add @niyongsheng/free-vision-skill
Then add to cordis.patch.yml:
- insert:
- id: free-vision-skill
name: '@niyongsheng/free-vision-skill'
config:
timeout: 120000
Tools
view_image— describe image content (scene, people, QR, composition)ocr_image— extract text;layout=truefor table structure + coordinates
Input: http(s) URL / base64 / local path.
Paste-to-path (Web UI)
Paste (⌘V) an image in the DSH web input box → its local absolute path is inserted. Loopback-only upload, magic-byte checked: PNG / JPEG / GIF / WebP / HEIC / HEIF.
Usage (Claude Code Skill)
swift scripts/ocr.swift image.png # OCR
swift scripts/ocr.swift --layout image.png # table + coordinates
swift scripts/ocr.swift --describe image.png # describe image
Notes
- Requires macOS 11+ & Xcode Command Line Tools
- First run compiles ~5–10s, cached afterwards
License
MIT © 2026 Nico
链接
同类插件
liustack/modlens★ 1837
为纯文本模型架起视觉桥梁:粘贴图片,输出结构化 JSON 证据(OCR、版面、语义)。
Anionex/dsh-vision-toolkit★ 422
让纯文本模型更好地做视觉任务:带意图的图片问答、长截图 OCR、UI 还原等。
superdesigndev/treg★ 416
给 Agent 的工具目录:按「要做的事」检索约 2,600 个外部接口(SEO 与 SERP、外链、社交、人物与公司信息补全、广告库、抓取),查看参数与单次调用价格后直接调用,凭据由服务端注入。附带技能,MCP 行在未设置 TREG_TOKEN 前保持禁用。
Lum1104/dsh-browser★ 156
Chrome 侧边栏扩展,让 DSH 直接操控你的浏览器,无需视觉能力。
zhaoolee/notes★ 142
将 DSH 对话导出为锤子便签风格 PNG,或在配置的账号工作区中新建和更新 Markdown 便签。
ysr666/dsh-vision-router★ 135
为纯文本 Agent 提供视觉能力:内置免 Key 视觉链 + 像素级视觉工具(看图问答、定位、裁剪、像素对比、取色、OCR、矢量化、抠图、截图);粘贴图片即可用。