DSH 的本地 Windows 11 OneOCR 工具:`oneocr_recognize` 返回 OCR 文本,以及包含行/词多边形、置信度、旋转角度和手写体样式的结构化结果。
安装
# GitHub 源码(首次需按提示配置 allowBuilds 构建授权后重试)
dsh plugin --profile web add github:hawkhai/win11-oneocr
装任何插件都等于在你的机器上跑第三方代码,权限和你本人一样大——能读你的文件、用你的凭据、访问网络,工具审批管不到它。GitHub 来源的插件还会在安装时执行构建脚本——pnpm 默认拦截,所以安装可能停在 ERR_PNPM_GIT_DEP_PREPARE_NOT_ALLOWED 或 ERR_PNPM_IGNORED_BUILDS;dsh 会打印出需要添加的确切键名,把它加进该 profile 的 pnpm-workspace.yaml 的 allowBuilds 下,重跑一次即可装上。放行构建本身就是一次信任判断:请只安装可信来源,并尽量锁定 commit(github:owner/repo#sha)。
README
该插件的 README 只有英文版本。
Offline OCR engine extracted from the Windows 11 Snipping Tool, with full-featured C++ CLI, reusable DLL wrapper, and Python visualization.
Based on: https://b1tg.github.io/post/win11-oneocr/
DeepSeek Harness plugin
This repository can be installed as a DeepSeek Harness plugin:
dsh plugin add github:hawkhai/win11-oneocr
It registers oneocr_recognize, a model-facing tool that accepts a local image path and returns recognized text together with OneOCR's structured line/word polygons, confidence values, image angle, and handwriting style. The tool runs locally on Windows 11; image bytes are not sent to an external OCR service.
The bundle defaults to the prebuilt bin/ocr.exe and its adjacent runtime files. Override ocrBin, timeoutMs, or maxOutputBytes in the plugin row if needed.
Features
| Feature | Description |
|---|---|
| 8-point Bounding Box | 4-corner polygon bbox for lines and words (not just axis-aligned rect) |
| Word Confidence | Per-word recognition confidence score (0.0–1.0) |
| Image Angle | Detected rotation angle of the text in the image |
| Line Style | Handwritten vs. printed text classification with confidence |
| Resize Resolution | Configurable max internal resize before OCR (performance/accuracy trade-off) |
| Resource Release | Proper cleanup via ReleaseOcrResult, ReleaseOcrPipeline, etc. |
| Unicode Path | Full Unicode file path support via _wfopen in the DLL wrapper |
| Multi-image Batch | Process multiple images in one invocation |
| Plain Text Output | --text mode for pipe-friendly output (no JSON) |
| Raw Buffer OCR | ocrImageRaw() for in-memory BGRA pixel buffers (no file I/O) |
| Visualization | Python script with confidence-colored word boxes and style labels |
Prerequisites
- Windows 11 (tested on 23H2+)
- Snipping Tool 11.2409.25.0+
Copy these 3 files from the Snipping Tool installation folder into the same directory as ocr.exe:
oneocr.dlloneocr.onemodelonnxruntime.dll
Find the Snipping Tool folder:
Get-AppxPackage Microsoft.ScreenSketch | Select-Object -ExpandProperty InstallLocation
Example: C:\Program Files\WindowsApps\Microsoft.ScreenSketch_11.2409.25.0_x64__8wekyb3d8bbwe\SnippingTool
CLI Usage (ocr.exe)
ocr.exe <image1.png> [image2.jpg ...] [options]
Options
| Option | Description |
|---|---|
--text, -t |
Output plain text only (no JSON) |
--output, -o <file> |
Write JSON to specified file (default: <image>.json) |
--max-lines <n> |
Max recognition lines, 1–1000 (default 1000) |
--resize <WxH> |
Max internal resize resolution (e.g. 1152x768) |
--quiet, -q |
Suppress progress messages |
--help, -h |
Show help |
Examples
# Single image → JSON
ocr.exe screenshot.png
# Plain text output (pipe to file)
ocr.exe screenshot.png --text > result.txt
# Batch process
ocr.exe img1.png img2.jpg img3.bmp
# Custom options
ocr.exe photo.jpg --max-lines 50 --resize 800x600 -o result.json
JSON Output Format
{
"file": "test.png",
"image": { "width": 771, "height": 479, "step": 3084 },
"image_angle": 0.0643,
"line_count": 2,
"lines": [
{
"index": 0,
"text": "Hello World",
"bounding_box": [
13.0, 38.0, 458.0, 38.0,
458.0, 77.0, 13.0, 76.0
],
"style": { "type": "printed", "confidence": 0.035 },
"word_count": 2,
"words": [
{
"index": 0,
"text": "Hello",
"bounding_box": [
14.35, 39.70, 140.35, 41.31,
139.93, 73.42, 13.78, 74.09
],
"confidence": 0.987
}
]
}
]
}
DLL Wrapper (oneocr_wrapper.dll)
A reusable C DLL wrapper with 3 main APIs:
| Function | Description |
|---|---|
initModel(model_dir) |
Load DLL + model, initialize pipeline |
ocrImage(image_path, json, alloc) |
OCR an image file → JSON string |
ocrImageEx(image_path, json, alloc, max_lines, resize_w, resize_h) |
OCR with configurable options |
ocrImageRaw(pixel_data, w, h, step, json, alloc) |
OCR on raw BGRA pixel buffer |
releaseModel() |
Clean up all resources |
C++ Header-Only Usage (oneocr.h)
#include "oneocr.h"
OneOcr ocr; // loads oneocr_wrapper.dll
ocr.initModel(L"."); // directory with oneocr.dll + .onemodel
std::string json;
ocr.ocrImage(L"test.png", json); // basic OCR
ocr.ocrImageEx(L"test.png", json, 50); // max 50 lines
ocr.ocrImageRaw(bgra_ptr, w, h, json); // raw buffer OCR
Visualization (visualize.py)
python visualize.py <image_path> <json_path> [output_path]
Features:
- 8-point polygon bounding boxes (lines in red, words colored by confidence)
- Confidence score labels below each word
- Handwritten lines highlighted in orange, printed in red
- Image angle and line count overlay
Build
Requires: MSVC (Visual Studio), json.hpp (nlohmann/json), stb_image.h.
# Build CLI
cl /EHsc /O2 ocr.cpp /Fe:ocr.exe
# Build wrapper DLL
cl /EHsc /O2 /LD oneocr_wrapper.cpp /Fe:oneocr_wrapper.dll
# Build test
cl /EHsc /O2 oneocr_test.cpp /Fe:oneocr_test.exe
Other Implementations
| Directory | Language | Description |
|---|---|---|
oneocr/ |
Python | PyPI package with PIL/cv2 input, FastAPI web server |
oneocr-rs/ |
Rust | crates.io library with image crate, serde JSON |
oneocr-cli/ |
Rust | Minimal CLI, plain text output |
win11_oneocr_py/ |
Python | Basic ctypes script (original Python port) |
Credits
- b1tg - Original reverse engineering and C++ implementation
- AuroraWright/oneocr - Python package with web server
- wangfu91/oneocr-rs - Rust binding with full feature coverage
- Cecilia-pj/win11_oneocr_py - Original Python port
License
MIT
链接
同类插件
liustack/modlens★ 4103
为纯文本模型架起视觉桥梁:粘贴图片,输出结构化 JSON 证据(OCR、版面、语义)。
ysr666/dsh-vision-router★ 1127
为纯文本 Agent 提供视觉能力:内置免 Key 视觉链 + 像素级视觉工具(看图问答、定位、裁剪、像素对比、取色、OCR、矢量化、抠图、截图);粘贴图片即可用。
Anionex/dsh-vision-toolkit★ 887
让纯文本模型处理视觉任务:粘贴图片后自动切换到 Vision Toolkit 变体,支持图片问答、多图比较、长截图 OCR、截图还原前端 UI、元素定位与像素对比。默认无需 API Key——图片经作者自建的免费服务处理,每台机器每天 100 张;也可改为指向自己的服务商。
dickpy/dsh-imagegen★ 99
面向 DSH Web GUI 的 AI 生图插件:通过可配置的 OpenAI 兼容端点(gpt-image-2 / gpt-image-1 / dall-e-3)实现文生图与图生图,提供 api_url/api_key 设置卡片与侧边栏分栏生图工作台。
fandc520/dsh-comfyui★ 93
让 DeepSeek Harness 的 Agent 直接驱动本地或远程 ComfyUI:comfyui_run / comfyui_object_info / comfyui_workflow 工具生成与编辑图像、视频,附带工作流库(图工作流提取:按分量 / 主流程 / 整体)、加载区分辨率自动匹配、实时队列、SDXL 与 Wan 2.1 模板、配套 skill 与同源媒体代理。
sunxin-ai/dsh-design-qa★ 44
给纯文本模型的设计稿保真判定:`deepseek_vision` 工具从任意 OpenAI 兼容视觉路由借来一只眼,让模型判断实现与设计稿是否一致——并附上支撑该判定的基准(4 组夹具、23 处注入缺陷、逐格原始输出)与其依赖的提问纪律。
社区评论
评论公开保存在 GitHub Discussions。加载评论会连接 GitHub 和 Giscus;发表内容需要 GitHub 账号。