PDF parsing tools for DSH: pdf_info, pdf_extract_text, and pdf_render_page with dual pdfjs and built-in rendering engines, including system-font rendering for PDFs with non-embedded CJK fonts.
Install
# from npm (prebuilt)
dsh plugin --profile web add @zhtx2026/dsh-pdf
# from GitHub (first run asks for allowBuilds approval — follow the hint, retry)
dsh plugin --profile web add github:zhtx2024/dsh-pdf
Any plugin you install runs third-party code with your own permissions — it can read your files, use your credentials, and reach the network, and tool approvals don’t sandbox it. GitHub-sourced plugins also run build scripts at install time — pnpm blocks those until you allow them, so an install can stop with ERR_PNPM_GIT_DEP_PREPARE_NOT_ALLOWED or ERR_PNPM_IGNORED_BUILDS; dsh prints the exact key to add under allowBuilds in your profile’s pnpm-workspace.yaml, and the install works on the next run. Allowing a build is a trust decision: only install sources you trust, and pin a commit (github:owner/repo#sha).
README
PDF parsing toolkit for DeepSeek Harness — three agent-facing tools (
pdf_info/pdf_extract_text/pdf_render_page) with hybrid engines, including system-font rendering for PDFs with non-embedded CJK fonts (e.g. EasyEDA / JLC EDA schematic exports).
Tags: dsh-plugin · pdf · pdfjs · canvas · schematic · tools
Install
dsh plugin --profile web add @zhtx2026/dsh-pdf
Or from a local checkout during development:
dsh plugin --profile web add link:<path-to-repo>
Then restart DSH (or open a new session) — the three tools and an agent guidance block become available.
Tools
| Tool | What it does |
|---|---|
pdf_info |
Page count, page sizes (pt), document metadata, and the font list with per-font embedded flags. |
pdf_extract_text |
Extracts text per page; optional page and withPosition (per-line x/y coordinates + font size). |
pdf_render_page |
Renders one page to a PNG (local file) with configurable scale/outDir; engine auto-selection. |
Why a custom renderer?
pdfjs-dist cannot render text in PDFs whose fonts are not embedded and lack a
ToUnicode map — the classic case is JLC EDA / EasyEDA schematic exports
(SimSun/SimHei Type0 fonts with UniGB-UCS2-H encoding): pdfjs emits
translateFont failed and drops the glyphs.
dsh-pdf solves this with a two-engine design:
- pdfjs engine — used when all fonts are embedded (standard PDFs).
- sysfont engine — a built-in content-stream renderer that decodes the
PDF text operators (
BT/ET,Tf,Tm,Td,Tj,TJ) itself, mapsUniGB-UCS2-H/UTF-16BE strings to Unicode, and draws them with system fonts (SimSun,SimHei,Microsoft YaHei, …) via@napi-rs/canvas.
Text extraction uses the same auto-selection: embedded fonts → pdfjs text layer; non-embedded fonts → the built-in decoder (which correctly recovers CJK text pdfjs loses).
Example
pdf_info("D:/Downloads/SCH_SA-V11A.pdf")
→ 3 pages, A3 landscape, jsPDF producer, 6 fonts (all non-embedded)
pdf_extract_text("D:/Downloads/SCH_SA-V11A.pdf", { page: 1 })
→ net names, pin numbers, and Chinese labels ("主控板", "嘉立创", …)
pdf_render_page("D:/Downloads/SCH_SA-V11A.pdf", { page: 1, scale: 2 })
→ D:/Downloads/SCH_SA-V11A-pages/page-1.png (2396×1698, engine: sysfont)
Architecture
lib/index.js— cordis plugin entry (name,inject,apply); registers the three tools as native tool objects (plain JSON-Schemaparametersandoutput— no@deepseek-ai/dsh-toolsdependency needed) and injects the agent guidance block (systemPrompt.section, order 160).lib/pdf-core.js— self-contained engine: PDF object parser, font scanner, pdfjs loading with fs-basedCMapReaderFactory/StandardFontDataFactory(Node 24's globalfetchdoesn't speakfile://), the built-in text extractor, and the sysfont renderer.cordis.patch.yml— thedsh.bundle.patchinsert row.
Limitations
- Scanned/image-only PDFs have no text layer — extraction returns nothing (rendering still works, as a plain image).
- The sysfont renderer assumes
Type0CID codes equal Unicode code points (true forUniGB-UCS2-Hexports); exotic CID fonts fall back to pdfjs. - Large PDFs are parsed fully into memory.
License
Links
More in this category
tt-a1i/archify#integrations/deepseek-harness★ 74360
Generate validated, self-contained interactive architecture, workflow, sequence, data-flow, and lifecycle diagrams from repositories or system descriptions.
dream-num/dsh-univer-office★ 442
Give DeepSeek Harness a real office environment. Univer Office Plugin brings spreadsheets, docs, slides, canvases, relational tables, and more into one runtime — with connected data, validation, versioned changes, and isolated worktrees for multi-agent collaboration.
PerryLink/dsh-industry-research★ 186
Deterministic industry research reports for DeepSeek Harness — company and industry research flows produce structured, verifiable reports from staged evidence.
HuanLinOTO/dsh-plugin-mineru★ 47
Expose MineRU document parsing tools to the model.
kw78/dsh-office-tools★ 26
Workspace-safe Office tools for agents: create/read Word, create/read/update Excel, and create/read PowerPoint decks with PNG/JPG/GIF image placement.
PolinniZhong/dsh-knit★ 18
Lists the Markdown documents, images and video that already exist anywhere in the session workspace in the DSH sidebar, ranked by relevance to the current conversation: recent messages are matched locally against document title, summary and body with IDF weighting, with no model calls and no network. Because the list is scanned from the workspace instead of remembered, restarting DSH or starting a new session does not empty it. Images and video preview in place, with relative-path images resolved and video streamed over HTTP Range. A references bar under the preview header shows which documents cite the one being previewed and which it cites, with one click to jump between them. The same ranking is exposed to the agent as a knit_docs tool, which returns the most relevant documents along with the passage that matched in each, where one is found.
Community comments
Comments are public GitHub Discussions. Loading them connects to GitHub and Giscus; a GitHub account is required to post.