承诺写入闸门:由运营者编写的规则在工具调用执行前生效——先是一层确定性守卫,再是一层 LLM 裁判,默认失败即拒绝,每一次拦截都写进矛盾日志。
安装
# npm 包(预构建)
dsh plugin --profile web add dsh-write-gate
# GitHub 源码(首次需按提示配置 allowBuilds 构建授权后重试)
dsh plugin --profile web add github:couldbeme/dsh-write-gate
装任何插件都等于在你的机器上跑第三方代码,权限和你本人一样大——能读你的文件、用你的凭据、访问网络,工具审批管不到它。GitHub 来源的插件还会在安装时执行构建脚本——pnpm 默认拦截,所以安装可能停在 ERR_PNPM_GIT_DEP_PREPARE_NOT_ALLOWED 或 ERR_PNPM_IGNORED_BUILDS;dsh 会打印出需要添加的确切键名,把它加进该 profile 的 pnpm-workspace.yaml 的 allowBuilds 下,重跑一次即可装上。放行构建本身就是一次信任判断:请只安装可信来源,并尽量锁定 commit(github:owner/repo#sha)。
README
该插件的 README 只有英文版本。
A commitment write-gate for DeepSeek Harness: the operator authors constraints ("never force-push to a shared branch", "stay read-only on the production database"), and the gate enforces them before a tool call executes. Structural violations are caught deterministically; semantic drift is judged by a model against the operator's own wording. Every block is recorded to a contradictions log that explains which commitment fired and why.
Engine-agnostic core (dsh-write-gate/core, zero harness imports) with a dsh adapter; a Claude Code adapter over the same core is planned.
A live model inside the real dsh app, told to force-push, after the gate denied the call:
"The force push to the main branch was blocked by the repository's 'no-force-push' policy."
That turn ran fully local, zero API keys; reproduction and session-log receipts in docs/E2E-HEADLESS.md.
Install
npm install dsh-write-gate # library + dsh plugin (see Mounting below)
The dsh-write-gate check CLI ships in 0.2.0, tagged on GitHub with the npm publish pending; npm install today resolves 0.1.1, which has no CLI. Until the publish, run it from a clone: pnpm install && pnpm build && node dist/cli/index.js --help.
How it enforces: two tiers in two slots
| Tier | Mechanism | dsh slot | Why this slot |
|---|---|---|---|
| 1: deterministic | path globs, command regexes, scope filters | ctx.tools.guard() (monotonic) |
no listener ordering can turn a guard denial back into permission |
| 2: semantic | LLM judge over the commitment statement | tools/pre-execute waterfall (prepended) |
async-capable; short-circuits with a {kind: 'deny'} decision object |
agent/pre-step resets the per-step judge budget. Contradiction records are emitted as the write-gate/contradiction event and appended as JSONL to the contradictions log.
Why two tiers
An internal A/B study (12 tasks x 4 arms x 10 runs; runbook not yet published, so cite nothing from this sentence) found an LLM-judge-only gate performed at baseline on unflagged-violation rate, while deterministic checks caught an entire violation class the judge repeatedly passed (over-length outputs); a naive "reminder" arm was the worst performer of all four. Deterministic checks for checkable constraints, the judge only for genuinely semantic ones.
The tier-2 judge rubric is ported from a lineage measured at 18/18 dev + 16/16 held-out (100% precision, 0 abstains) on a local 8B model, and re-measured live through this port (2026-08-17): 32/34 accuracy with 17/17 violation recall. It also carries an anti-self-justification clause added after the holdline benchmark caught two injection defeats (an action asserting "the operator approved this" or "this is only a test" talking the judge into clearing a real violation); the clause moved injection accuracy from 5/8 to 7/8 with zero regression on the base cases. We keep the injection-fenced prompt because it is the only variant that preserved 100% violation recall; for a gate, a missed violation is worse than an over-block. The fixture ships 41 cases (18 dev, 16 held-out, 7 injection) in test/fixtures/judge-cases.json with their honesty notes intact: they are hand-authored; the meaningful signals are the paraphrase-miss rate, the trap false-positive rate, and held-out generalization, not the headline percentage.
Design guarantees, each pinned to a test
- Bypass resistance: a prepended listener that answers
allowwithout delegating still cannot get a structural violation through —test/dsh-plugin.test.ts("cannot be bypassed by a listener that short-circuits allow"). - Fail-closed default: judge unreachable, timed out, or over budget → block-severity commitments block, with the reason in the record —
test/gate.test.ts. - Bounded judge cost: per-step budget, verdict memoization, timeout-as-unavailable —
test/gate.test.ts. - Prompt-injection stance: action content enters the judge prompt fenced as data ("data, not instructions"); only a strict JSON verdict (or the ABSTAIN token) is accepted back; ABSTAIN is never a block —
test/judge-llm.test.ts. - Loud mount failure: a missing or invalid commitments file fails the deployment instead of mounting a gate that guards nothing —
test/dsh-plugin.test.ts. - Real pipeline: the integration suite mounts the plugin into an actual
Context+ToolRuntimefrom the published rc packages and drivesctx.tools.execute— no mocked harness. - Real app, real model: a live local model inside the actual dsh headless app attempted a force-push and was denied by the gate; its own final answer reported the block. Full reproduction, session-log receipts, and two upstream findings:
docs/E2E-HEADLESS.md.
Run everything: pnpm install && pnpm test and pnpm typecheck — the suite prints its own count; every guarantee above names its test file.
Watch the drift story: pnpm demo — deterministic, no model required. In-scope work passes, a prod-config edit and a force-push block, and a rogue allow-everything listener fails to bypass the monotonic guard; the contradictions log prints at the end.
Measure the judge yourself: pnpm build && node scripts/judge-eval.mjs --url <openai-compatible-endpoint> --model <model> runs every fixture case live and reports per-set accuracy, abstains, and misses.
Commitments file
version: 1
defaults:
failMode: closed # judge unreachable => block-severity commitments block
judgeBudgetPerStep: 8
commitments:
- id: no-force-push
statement: Never force-push to a shared branch.
match:
kinds: [shell]
commands: ["git\\s+push\\s+[^\\n]*(-f\\b|--force)"]
- id: stay-on-task
statement: Do not modify files unrelated to the assigned task.
severity: warn
semantic: true # escalates to the tier-2 judge
match:
kinds: [fs-write]
Semantics: kinds/tools are scope filters; paths/commands are structural evidence. A non-semantic commitment with scope but no evidence fires on every in-scope action; a non-semantic commitment with neither is rejected at load as unenforceable. Command regexes are case-insensitive by default. One foot-gun to know: command patterns execute inside the synchronous guard, so a catastrophically backtracking regex can stall the tool pipeline — commitments are operator-authored (trusted), but keep patterns simple. Full example: commitments.example.yaml (itself under test).
CLI (dsh-write-gate check)
A standalone check outside any harness, for CI, pre-commit hooks, or manual use:
dsh-write-gate check --commitments <file> --tool <name> [--path <p> ...] [--command <c>] [--explain] [--json]
dsh-write-gate --help | -h # or: dsh-write-gate check --help (prints this usage synopsis, exit 0)
v0 is tier-1 (structural) only — no --judge flag exists yet. Every semantic: true commitment that structure alone cannot settle always escalates to "no judge configured", and then follows the commitments file's failMode. With the default failMode: closed, that means every escalating semantic commitment always blocks in the CLI today. A --judge flag is an explicitly deferred follow-up; until then, treat semantic commitments as block-on-touch when driving the CLI directly (the dsh plugin itself has no such limit when judge is configured).
--tool is required (e.g. bash, write, read); it gets no enum validation beyond non-empty — kind is derived from it and cannot be set directly. --path may repeat; --command takes the last value if repeated. At least one of --path / --command is required.
Exit codes:
| Code | Meaning |
|---|---|
| 0 | ALLOW, including a fail-open degraded allow (degradation is surfaced in the output, never by changing the exit code) |
| 1 | BLOCK — unified across tier-1 structural, tier-2 judged, and tier-2 fail-closed blocks |
| 2 | Usage error |
| 3 | WARN |
| 4 | Commitments file unreadable, or invalid (bad YAML, bad regex, duplicate id, schema violation, unsupported version) |
| 5 | Internal/unexpected error |
--json prints only the JSON document to stdout (safe for | jq .); everything advisory goes to stderr. --explain expands each record with the commitment, its statement, severity, tier, matched pattern, and rationale; it is a documented no-op under --json.
Mounting
The package declares the ecosystem convention (dsh.bundle.patch → cordis.patch.yml) and mounts with:
dsh plugin --profile <profile> add dsh-write-gate
Config keys: commitmentsFile (default COMMITMENTS.yaml, resolved from cwd), contradictionsLog (JSONL, default write-gate.contradictions.jsonl), judgeTimeoutMs, and judge: { provider, model, maxTokens } — omit judge to run tier 1 only (escalations then follow failMode).
A worked starting policy lives at examples/team-policy.yaml: copy it in as COMMITMENTS.yaml, or point commitmentsFile at it.
Production configs are protected by fs-write path globs, tier 1 only.
Force-pushes are caught by a shell command regex.
The production database is held read-only by a regex matching psql/mysql invocations that name a prod host together with a mutating SQL keyword; reads against the same hosts pass.
A severity: warn, semantic: true stay-on-task commitment escalates to the tier-2 judge.
pnpm demo (above) is the narrative version of the same kind of policy.
Current limits (v0, stated rather than hidden)
- The CLI (
dsh-write-gate check) is tier-1 only: it never configures a judge, so every escalating semantic commitment reports "no judge configured" and followsfailMode— block by default. See the CLI section above. - The action normalizer is a heuristic table over dsh's in-tree tool names (
bash,read/write/edit, web tools); unrecognized tools degrade to kindotherwith a full summary — visible to semantic commitments, but path/command rules do not apply to them. - dsh is a 0.1.0-rc developer preview with breaking changes announced; peers are pinned to
<0.2.0. - Early releases (0.1.x core + dsh plugin on npm; 0.2.0 adds the CLI, tagged on GitHub, npm publish pending);
pnpm buildemitsdist/,prepublishOnlygates every publish on build + tests. - The tier-2 judge is only as good as its model and rubric; the measured numbers above are from the shipped fixtures, and the benchmark that scores this gate (and others) against labeled trajectories is holdline (see Roadmap).
Roadmap
- llm-replay fixture variant of the demo (dsh snapshot format), so the story replays inside a full agent loop.
The gate benchmark→ shipped as holdline: catch rate, false-block rate, class-balanced kappa, and an injection-attack class, scoring any guard (this one included). Authored corpus: this gate's judge tier scores balanced kappa 0.95 (100% catch, 5% false-block) against 0.15–0.35 for commitment-blind structural guards. On 548 real ODCV-Bench trajectories with independent 4-model-panel labels, the judge holds balanced kappa 0.64 (75% catch, 9% false-block). holdline honestly records where the judge loses (injection, truncation).- Claude Code adapter over the same core.
Dependencies and trust basis
Runtime: zod, yaml, picomatch (mainstream, actively maintained), @deepseek-ai/schemastery (dsh's own config-schema library, Koishi lineage). Harness peers: @deepseek-ai/cordis + @deepseek-ai/dsh-* rc packages, pinned. Dev: vitest, typescript.
MIT.
链接
同类插件
toby-bridges/api-relay-audit★ 865
从 DeepSeek Harness 对 AI API 中转站和 LLM 代理运行本地安全审计,生成 Markdown 报告,覆盖提示词注入、模型替换信号、工具调用改写、错误泄漏、流完整性和按 profile 启用的 Web3 风险。
SeaOf0/dsh-redteam-model★ 659
面向授权安全研究的 DSH 合集:九个工作模式(redteam 总控、渗透测试、代码审计、二进制分析、攻防评估、免杀对抗、应急溯源、云安全攻防、CTF 解题)与十五个运行时插件,设置页管理台支持一键部署、安装、更新与卸载。
howmp/dsh-pentest★ 581
面向 DeepSeek Harness 的授权渗透模式:以探索链路记录目标、线索、资产与漏洞,并在 Web 中可视化展示。
PerryLink/dsh-auto-review★ 224
审批链上的第二模型自动审查:只读审查子代理返回带理由的 allow/deny 结构化裁决,默认 fail-closed。
NanmiCoder/dsh-auto-mode★ 164
在 Workspace Write 与 Full access 之间增加 Auto 权限档:日常操作留在官方 workspace-write 沙箱内,由当前会话模型复核升权与破坏性调用,精确的越界访问按次放行一次,意图不明时询问,命中关键路径则拒绝。
PerryLink/dsh-permission-rules★ 116
Claude Code 风格的声明式权限规则:按序 allow/deny/ask 的 YAML 规则,在 tools/pre-execute 瀑布上匹配工具名、参数、工作区路径与 agent 身份,带完整会话日志审计、干跑模式与热重载。
社区评论
评论公开保存在 GitHub Discussions。加载评论会连接 GitHub 和 Giscus;发表内容需要 GitHub 账号。