Evidence-first guard: warns the agent when it claims success without a matching tool-execution record in recent context.
Install
# from GitHub (first run asks for allowBuilds approval — follow the hint, retry)
dsh plugin --profile web add github:jinguanghai/deepseek-harness-forge-plugins#path:/plugins/evidence-first
Any plugin you install runs third-party code with your own permissions — it can read your files, use your credentials, and reach the network, and tool approvals don’t sandbox it. GitHub-sourced plugins also run build scripts at install time — pnpm blocks those until you allow them, so an install can stop with ERR_PNPM_GIT_DEP_PREPARE_NOT_ALLOWED or ERR_PNPM_IGNORED_BUILDS; dsh prints the exact key to add under allowBuilds in your profile’s pnpm-workspace.yaml, and the install works on the next run. Allowing a build is a trust decision: only install sources you trust, and pin a commit (github:owner/repo#sha).
README
Evidence-first guard plugin for DeepSeek Harness. 证据铁律:声称完成必须有实际执行证据。
Why
LLMs are statistical machines: they can hallucinate, misattribute, and — most dangerously — claim completion without actually executing anything. Tool architectures can structurally prevent phantom tools and phantom calls (schema validation, closed registries), but no architecture can prevent the fourth kind of hallucination: a model that says "done" with no evidence behind it.
This plugin closes that gap at the session layer:
- Records every tool execution (
tool/resultevents) per turn. - Scans assistant messages for completion claims (完成/成功/修复/搞定/通过…).
- A claim with no nearby tool execution → the next
pre-stepinjects a visible warning, forcing the model to either supply evidence or retract. evidence_audittool: full audit report for the human gatekeeper.
Install
# via cordis.patch.yml / bundle
- id: evidence-first
src: link:./plugins/evidence-first
Or copy lib/index.js into your plugin directory and register it.
Config
| Field | Type | Default | Meaning |
|---|---|---|---|
claimPatterns |
string[] |
Chinese completion phrases | Regex sources for completion claims |
evidenceWindowTurns |
number |
1 |
How many turns back counts as "nearby evidence" |
injectWarnings |
boolean |
true |
Inject visible warnings on the next pre-step |
registerAuditTool |
boolean |
true |
Register the evidence_audit tool |
maxEvidenceEntries |
number |
200 |
Per-session evidence cap (memory guard) |
Config is fail-loud: invalid values throw at load time, never silently fall back.
How it works
session/event— observeturn/start,tool/result,assistant/messageagent/pre-step— inject the pending warning into the next model requestctx.tools.register—evidence_auditaudit tool- Injected messages carry
source: { kind: 'plugin', plugin: 'evidence-first', form: 'notice', summary: '证据铁律警告' }— the{kind:'plugin'}tag is load-bearing: untagged context would render as a user prompt.
Design notes
- Prefers false positives over false negatives. A warning is cheap; a silent unverified "done" is expensive. Claims are flag-for-human, not auto-rejected.
- Zero runtime dependencies. The plugin imports nothing from
@deepseek-ai/*at runtime — it compiles standalone and loads in any dsh environment. - Evidence window defaults to 1 turn (same turn or the previous turn).
Adjust via
evidenceWindowTurnsfor longer tool chains.
License
MIT
✅ Official convention compliance (官方规范合规)
DeepSeek Harness CONTRIBUTING.zh.md
states that the project cannot accept external PRs, and directs the community
to create and share plugins (tag repos with dsh-plugin). This plugin follows
that path — it is a drop-in Cordis plugin, built to the same conventions as
first-party plugins:
| Official convention | This plugin |
|---|---|
Named exports { name, Config, inject, apply } |
✅ same pattern (apply is default; inject injects the evidence contract into the model prompt) |
Config with fail-loud validation (issues array + throw in apply) |
✅ invalid patterns → issues; semantic errors → loud failure |
Assembly via cordis.patch.yml bundle |
✅ one-layer insert patch |
| Zero modification of official source | ✅ pure event hooks (turn/start, assistant/message, tool/result, pre-step) + one registered tool (evidence_audit) |
package.json per first-party standard |
✅ exports/types/files/peerDependencies (@deepseek-ai/cordis, dsh-agent, dsh-tools) |
i18n README (README.i18n.yaml) |
✅ en/zh |
| Zero npm runtime deps | ✅ except @deepseek-ai/schemastery (Config validation, same as first-party) |
Verification (evidence first, of course): 5 unit tests (warning injection on evidence-less claims, no false-positive with real tool execution, tool registration, fail-loud Config, source-tag integrity) — 5/5 pass; plus an end-to-end headless run where the model actually executed a shell tool and claimed completion → no false warning, no inject errors.
Share: repo jinguanghai/deepseek-harness-forge-plugins (tagged dsh-plugin).
Links
More in this category
yjh051108/dsh-routing-suite★ 6995
One repository, three parts: a runtime injector for DSH plugin packages (inject, hot-reload, unload, promote a dev staging tool to the front, route self-heal, plus a settings-page plugin manager that lists, unloads and drags folders in to internalize), a task-aware reasoning-mode router agent preset (router-standard / router-spec / router-react), and a graded two-level task protocol whose six tools (commit_star, lock_stage, revise_do, edit_plan, mark_task, redteam_verdict) pin task state to disk. The injector implementation ships in-tree, so the install carries its own behaviour rather than a dependency list.
strukto-ai/mirage#dsh★ 3667
Swaps the filesystem and bash providers for a mirage virtual workspace: file tools and shell commands run over mounted resources (RAM, S3, Redis, Slack, Gmail, Notion, Postgres) instead of the host disk, with per-mount read/write/exec modes, per-command sandbox routing (monty, pyodide, quickjs in process; docker, e2b, daytona remote), and installed CLIs (git, gh, slack, linear, ntn, gws, or one you register) as head words in the virtual terminal.
hust-open-atom-club/oh-dsh★ 325
Community distribution: TUI, desktop, and Web UI as one bundle with layered installation.
weijiafu14/pi2dsh★ 208
Pi Host ABI compatibility engine: after one install, unmodified Pi extensions from npm mount as native DSH plugins with `dsh plugin add <pi-package>`. Verified end to end on stock DSH with pi-mcp-adapter (full MCP manager: OAuth, resources, prompts, MCP Apps, elicitation, sampling), @tintinweb/pi-subagents, pi-code, pi-hermes-memory and pi-background-tasks; `pi2dsh inspect` reports a package's compatibility before installing.
lire1131/dsh-undo-savepoint★ 167
Undo/redo & rollback system for DSH: every config change is auto-snapshotted; undo/redo/restore to any version from the WebUI or the offline CLI/GUI tools (works even when DSH fails to boot).
Fishquito7/dsh-skill-mcp-panel★ 158
Manages DSH skills and MCP servers from the web settings: skill cards with hot enable/disable, workspace scopes, groups, batch migration and drag-and-drop import, plus stdio/HTTP MCP CRUD with connection tests, secret redaction and the unified dsh-panel CLI.
Community comments
Comments are public GitHub Discussions. Loading them connects to GitHub and Giscus; a GitHub account is required to post.