Guarded context compaction for DeepSeek Harness (dsh): the LLM proposes, deterministic guards dispose — eager per-atom shrink (extract/summary/false under verbatim guards) + lazy reference-graph eviction (0-LLM) + byte-exact recall from an append-only log.
Install
# from npm (prebuilt)
dsh plugin --profile web add dsh-argp
# from GitHub (first run asks for allowBuilds approval — follow the hint, retry)
dsh plugin --profile web add github:yoza10635/dsh-argp
Any plugin you install runs third-party code with your own permissions — it can read your files, use your credentials, and reach the network, and tool approvals don’t sandbox it. GitHub-sourced plugins also run build scripts at install time — pnpm blocks those until you allow them, so an install can stop with ERR_PNPM_GIT_DEP_PREPARE_NOT_ALLOWED or ERR_PNPM_IGNORED_BUILDS; dsh prints the exact key to add under allowBuilds in your profile’s pnpm-workspace.yaml, and the install works on the next run. Allowing a build is a trust decision: only install sources you trust, and pin a commit (github:owner/repo#sha).
README
中文 | English
dsh-argp — Two-engine context compaction: guarded per-atom shrink + deterministic reference-graph eviction
dsh-argp is a third-party context compaction engine for DeepSeek Harness (dsh) in its two-engine form (npm default = both engines on (Stage-1 + Stage-2), mounted on install with no manual configuration; set peratom: false explicitly to fall back to 0-LLM graph eviction only — see "Install & mount"):
- Stage-1 per-atom shrink (eager, per turn) — at turn end, the turn's atoms are shrunk, not discarded: the model picks
extract(verbatim excerpt) /summary(abridged; every dropped token is itemized into an audit ledger) /false(keep original) per atom, and deterministic guards decide whether a proposal lands — anextractmissing even one load-bearing token is rejected whole. The LLM only proposes; it never destroys. - Stage-2 reference-graph eviction (lazy, three-tier trigger) — on the atom reference graph (deterministic A→R pairing edges + model-declared semantic cites edges), whole atoms are evicted in reverse topological order with zero LLM calls in the eviction phase, so the compression budget is honored exactly; triggering is a three-tier ladder: turn-start proactive / mid-turn pressure prune / post-clamp auto-continue (see "Three-tier trigger").
- The append-only log is the single source of truth — originals of everything shrunk or evicted stay in the log forever; two-tier recall
recall_summary/recall_detailbrings them back on demand (text blocks verbatim, hash-locked by tests; tool-call arguments are a JSON semantic-equivalent reconstruction, annotated, when the host stores them as an object; the fidelity guarantee covers structured load-bearing tokens verbatim, not prose-level wording; large nodes page via from/limit). The context is a rendered view of the log, not the history itself.
Why
Summarizer-based compaction (an LLM rewriting history) pays three prices:
- Lossy — exact tokens (paths, error codes, config values) die first in any rewrite, and are unrecoverable;
- Cache wipeout — rewritten history changes the system+prefix every turn, invalidating cross-turn KV/prefix caches from the change point on;
- Uncontrolled ratio — summary length is at the model's whim; the budget is never honored.
ARGP's answer: the LLM stays in the loop, but in chains — its output is always an untrusted input proposal, and guards adjudicate under verbatim discipline; forgetting is deterministic — 0-LLM graph rules guarantee convergence and budget; history is immutable — the append-only log carries every original, backed by the recall contract (never guess).
Measured (30-turn synthetic multi-turn coding task, four-arm comparison, spike 37): the only arm that achieves both 7/7 probe fidelity and lower cost than the active baseline is the dual-engine arm (A); the traditional summary baseline (D) is cheapest but scores 5/7 — it swallows exact strings and key gist. The selling point is not "cheapest" but "cheapest under fidelity" (A's full cost components ≤ baseline C; 3.39× more expensive than D — that gap is the price of fidelity, priced openly).
How it works (dual-engine pipeline, reverse-topological pruning, shadow-price contract, module map) — see ARCHITECTURE.md.
Core mechanics
Stage-1: PeratomCompressor (eager entropy reduction)
- Deterministic gating (
gate.ts): pure predicates decide whether a turn is compressible (long user messages / tool results over 512 chars / non-version-chain members); the LLM only decides how. - Per-atom decisions: a single LLM call returns
{seq, level, text}decisions —extract(verbatim subset; the guard hard-rejects missing tokens) /summary(allowed; every dropped token is itemized intosummaryDroppedfor audit) /false(explicitly keep the original); long user messages go through splitting (dialog verbatim transcription + remainder aggregated as U-info, gaps assigned to info). - No-op guard: full-copy proposals with ≤5% gain count as
false— no replace event is emitted (a zero-gain replacement wastes a surface generation for nothing). - Tail-only replacement: replacements must satisfy sourceEventSeqs ⊆ the current turn's range (out-of-range is a bug, locked by assertion); prefix-fingerprint tests prove request-prefix stability across N turns of per-turn compression — the lifeline of cache economy.
Stage-2: ArgpGraphEngine (lazy graph eviction, 0-LLM)
- Atomization + graph building: deterministic edges (assistant → its tool results, via
toolCallId) + semantic edges (the model outputs{"cites":[{"t":"prefix","l":"c|s|x"}]}per contract, four levels: critical/supporting/contextual/isolated). - Topological eviction: repeatedly evict in-degree-0 atoms (ordered by edge level → effective importance → last-reference turn); evicted atoms' outgoing edges vanish and downstream atoms unlock pass by pass; the closure lifecycle (ACTIVE→COMPLETED→PRUNABLE→PRUNED) retires finished task closures whole.
- Budget honored exactly: trigger line window = contextWindow×0.8 (default), retain target = window×0.2; the trigger timing is the three-tier ladder (see "Three-tier trigger" below); the degradation chain is lifecycle→force→fail (the summarize tier is a stub that always returns null — conservative option a of A6, not implemented;
enableSummarizedefaults to false), converging to budget or failing explicitly — measured 200K → 160K trigger → 32K retained, landing exactly.
Three-tier trigger (the Stage-2 trigger ladder)
Stage-2's "lazy" is not a single threshold check but a three-tier ladder (introduced in 1.4.0, revised in 1.5.0):
- Turn-start proactive (L1): the pressure check runs only at each turn's first pre-step (
step === 1), threshold = window (default contextWindow×0.8). The estimate includes this step's user message that has been claimed but not yet flushed to the surface — the host claims the new user message before dispatching pre-step, so turn start sees the exact content, not a prediction; a large paste at turn start is never missed. Turn start also runs the per-atom LLM pass (atom shrink) + graph eviction. - Mid-turn pressure prune (L1', 0-LLM):
midTurnPrunedefaults to true — whenstep > 1and pressure is reached, only 0-LLM graph eviction runs (no per-atom LLM pass — that is a 79s–3min block, not worth it mid-turn), withturnGuardrelaxed tomidTurnTurnGuard(default 0, allowing this turn's older A/R to be evicted;recencyGuardstill protects the newest nodes). The overage comes from this turn's tool results; the old guard (turnGuard=1) protected the whole turn — the root cause of 1.3.x's mid-turn prune being "nearly ineffective". The prune lands in that pre-step ⇒ the same step's request is already slimmed ⇒ the turn continues naturally. - Post-clamp auto-continue (L3): when the output is externally clamped (
finish=max-tokensandoutputTokens <the request'smaxTokens= the host/adapter shrank the output budget — the true capacity-pressure signal): if the turn is about to end (agent/turn-stoppinghook) ⇒ force-prune in place +steera continuation message, so the same turn keeps advancing (the user need not send "continue"); otherwise (the turn is still running) ⇒ force-prune at the next pre-step. The continuation text can be overridden bycontinuationNotice(empty string = prune without continuing); the cap isreactiveRetries(default 2, per "consecutive clamps" episode; from the 2nd on, recency/turn guards are relaxed; once exhausted, control returns to the overflow path).
Config knobs: midTurnPrune (default true) / midTurnTurnGuard (default 0) / continuationNotice (default built-in sentence) / reactiveRetries (default 2); 1.4.0's midTurnActive remains as a compatibility alias (true = mid-turn prune on with the default guard, the 1.3.x comparison tier; false = off).
Bridging and recall
- CiteDeclarer (per turn): the model declares cross-turn reference edges over its window, fed to Stage-2 through the
injectEdgeschannel — measured recall efficiency ≈ 2.6× the edgeless arm (zoom pinpoints the right atoms). - RecallZoom:
recall_summary(seq)(compressed state) /recall_detail(seq, from?, limit?)(log original: text blocks verbatim, tool-call arguments JSON semantic-equivalent and annotated when stored as an object; large nodes page via from/limit); 4:1 budgeting (summary budget = 4× detail) — over-budget calls return guidance text instead of a hard refusal. Atoms evicted from history are additionally served byrecall_pruned/list_pruned.
Model requirements (honest version)
The quality of per-atom split/shrink decisions depends on the model's instruction-following ability; the guards guarantee "never compresses badly" on any model (errors only ever err toward compressing less), but the benefit scales with compliance:
- Measured baseline: local Qwen3.6-35B-A3B / Qwen3.8-27B, full pipeline over 30/60 turns with 0 errors and 7/7 probes; parse failures of the split-transcription notation 0% (vs 72% for interval locating).
- Known DeepSeek-family trait: when the system prompt conflicts with user instructions (e.g. a task prompt saying "nothing else"), cites declarations can drop to 0 — semantic selectivity goes to zero, but Stage-1 guarded compression and Stage-2 deterministic eviction keep working and all invariants pass (50-turn v4-flash evidence). Once the task prompt leaves room for cites, the declaration rate recovers (43.6% measured over 10 turns).
- The information-contract compliance probe (info-contract) will be measured on DeepSeek in the post-1.0.0 recheck round.
Install & mount
Install from npm. npm default = both engines on: the package's bundle patch (cordis.patch.yml) declares maxPasses: 256 / recencyGuard: 10, and Stage-1's three pipelines are mounted by the engine at construction because peratom defaults to mounted (see below):
dsh plugin --profile <name> add dsh-argp
⚠️ Installing within 24 hours of a release: pin the version explicitly. Since pnpm 11,
minimumReleaseAgedefaults to 1440 minutes (24h): a just-published version is skipped and pnpm silently falls back to the newest version past the cooldown — no error, just the wrong version installed. After each dsh-argp release the older version's peerDeps usually no longer match the current host, so the host reportsPlugin dsh-argp@<older version> is incompatible with dsh <host version>. A version number in that error older than npmlatestmeans you hit this cooldown (unrelated to this repo or any mirror: npmmirror and npmjs agree ondist-tags).Two remedies (the first is narrowest):
dsh plugin --profile <name> add dsh-argp@<latest version> # ① pin the version, bypassing the cooldown# ② exempt this package in the profile's pnpm-workspace.yaml (durable; other packages stay protected) minimumReleaseAgeExclude: - dsh-argp
Disable the stock summarizer in the profile's cordis.patch.yml:
- id: compaction-basic
disabled: true
Mounting is handled by the package's own bundle patch (
cordis.patch.yml) (insertcreates the entry); the profile layer should only override config (modify) — do notinsertagain there (otherwiseduplicate loader entry id).
Inert under the minimal preset (since 1.9.0): ARGP stands down under the presets listed in
skipPresets— by default['minimal'], so sessions running the minimal preset are not pruned or compacted and their history stays pristine (as if this plugin were not installed). The decision follows the preset a session actually runs under (changeable via the preset-switch event); sessions with no preset info (headless/CLI) are unaffected. To make ARGP active under minimal too, addconfig: { skipPresets: [] }viamodifyin the profile layer.
Stage-1 (two-engine): mounted by default
The production mount path is the engine self-mounting from config.peratom at construction: peratom defaults to mounted — the three Stage-1 pipelines (compressor / declarer / zoom) are mounted and wired internally at construction, so a fresh install is already the two-engine form with no manual configuration.
To fall back to pure Stage-2 (0-LLM graph eviction), disable it explicitly:
- id: dsh-argp
config:
peratom: false # or null; both mean "Stage-2 only"
Why the default lives in code rather than in the package's bundle patch: patch-layer
configis replaced wholesale, not deep-merged (layer order bundle → profile → home → CLI). If the default were written only in the bundle patch'sinsertrow, any settings-page tweak (form values written back to the profile'scordis.patch.yml) would wipeperatomwhole and the two-engine form would silently fall back to Stage-2 only — the §11.13.1 gap reasserting itself.
Optionally pin a dedicated LLM backend for Stage-1 (otherwise it follows the host's routing):
- id: dsh-argp
config:
peratom:
compressor:
llm: { provider: deepseek-official, model: deepseek-v4-flash } # dsh-llm backend
declarer:
llm: { provider: deepseek-official, model: deepseek-v4-flash } # may point at a separate lite tier
# zoom: {} # two-tier recall; mounted by default when the block is present
The llm sub-block is optional: when both llm and endpoint/apiKey are absent, the backend automatically follows the host's routing (1.3.0's autoDshLlmSpec: in a real session it reads agent.options.{provider,model} on the fly + the host's ctx.llm service, following the host when it switches models); when all three paths fail to resolve, the component is naturally disabled (zero network). With llm set, it uses the host dsh-llm (production form); otherwise it falls back to OpenAI-compatible direct connection via endpoint/apiKey config or environment variables (the local llama.cpp experiment form). Stage-2 budgets are ratio-driven by default (window=ctx×0.8 / retain=window×0.2); no hardcoding needed.
presetClean is on by default: at mount time it generates a cleaned copy <id>-argp for each shipped preset that still carries the stock summarizer (removing compaction-basic/tool-result-pruner, keeping command-compact — /compact then routes to ARGP graph eviction); idempotent and fail-soft; set presetClean: false to disable it.
Validation results
Four-arm comparison (30-turn synthetic multi-turn coding, local Qwen3.6-35B-A3B, spike 37)
| Arm | Configuration | Probes | Cost (idle pricing) | Verdict |
|---|---|---|---|---|
| A dual-engine full | compressor + declarer + graph + zoom | 7/7 | ¥0.454 | The only 7/7 arm at ≤ baseline cost |
| B edgeless | declarer off | 7/7 | ¥0.556 | 31 recalls vs A's 12 (declarer ≈2.6× cheaper) |
| C active baseline | graph only (evicts at overflow) | 6/7 (R2 missed) | ¥0.802 | A's full cost components ≤ C |
| D summary baseline | dsh stock BasicCompactionEngine | 5/7 (loses exact tokens + gist) | ¥0.134 | Cheapest but loses fidelity — the foil for "cheapest under fidelity" |
Anti-interference: zero replacements of append-origin originals across arms A/B/C (A 140 / B 154 / C 216 events); prefix stability: A's 21 main requests share one fingerprint.
Water level and turn amplification
Measured behavior under a fixed window (16K tok; P5-bis, local Qwen3.6-35B-A3B, 2026-08-28 — evidence detail in CHANGELOG):
- Turn count amplified substantially: the zero-compaction control dies at ~20 turns on the window ceiling; the dual engine survives to full budget under the same window (~60-turn scale) — sustainable turns under a fixed window improve by an order of magnitude (lower-bound caliber; never cite bare multipliers — always carry window/task/model).
- Zero per-request prefix-stability degradation: the dual engine's per-request prefix fingerprint distribution matches the zero-compaction control; compaction events add only a one-time recompute tax, no cumulative degradation.
Graph engine historical validation (v0.3.x, DeepSeek v4-flash / Qwen3.8-27B)
50-turn t-long: U anchors 7/7, needles 7/7 (5/7 recovered via recall), 4 transactions 0 errors, compression target honored exactly (32K); 200K mainline cost ¥2.695 vs baseline high ¥3.087 (that baseline contains 77% empty-stream errors — a platform B-5 defect; read the contrast under this caliber, the disabled tier ¥3.19 is the cleaner control).
Reproduce
| Experiment | Command | What it validates |
|---|---|---|
| Four-arm comparison (needs a local model) | ARGP_ARM=A|B|C|D|E node spike/37-peratom-three-arm.ts |
Probe fidelity, cost triplet, anti-interference, K_no / amplification |
| per-atom soak | npm run spike36 |
Eight verdicts: gating / chain / conservation / prefix / VK-atom |
| 50-turn t-long | ARGP_DEEPSEEK_THINKING=enabled node spike/06-tlong.ts |
L1/L2/L3 invariants, anchors/needles, exact budget |
| Synthetic 0-LLM | npm run spike8a |
Single transaction with zero LLM calls |
| Per-atom audit | node spike/atom-audit.mjs <artifact dir> |
Event-driven per-atom shrink/eviction detail |
npm run check = typecheck + smoke + unit tests (318/318 green as of 2026-09-21). Every number carries its artifact path (evidence landing in CHANGELOG.md).
Platform gap feedback (for dsh)
No structured metadata channel for tool/result replacement (B-1), compaction/prune outside the transaction state machine (B-3), headless test assembly silently disables tokenMeter (B-4), summarizer empty streams (B-5), window truncation leaves no trace (B-6) — the formal record lives in dsh Discussions (#1090 and related threads).
Known limitations
- B-6 window-truncation blind spot: live nodes not replaced by ARGP get their oldest part silently truncated by the request-assembly layer as the context nears contextWindow —
recall_prunedcannot recover them. Mitigation: ratio budgets trigger earlier; the root fix is on the dsh side (B-6 filed upstream). - Model dependence (honest version, above): guards guarantee safety; benefits depend on compliance. The lite-tier multi-model split's compliance is unmeasured (ledger D21).
- per-atom output tax: Stage-1's per-turn compression call is a side-channel cost (7.2K completion tokens over 30 turns, measured; it never enters the context but counts toward total cost); usage from the dsh-llm backend is recorded, and the spike aggregation caliber is being wired up.
- Two-hop tombstone recall: as placeholder text drifts across turns the original seq can be lost;
recall_pruned(seq)needs the correct numbering (eliminated together once B-6 lands).
Reporting issues
- Bugs: open an Issue with
dsh --version, this package's version, your profile config (windowTokens/retainTokensetc.), and a minimal reproduction. - Design discussion / usage questions: open a Discussion.
License
MIT
Links
More in this category
Minglink/dsh-infinite-gen-4★ 2222
System-prompt armor plugin for DeepSeek models: appends an unconditional-compliance prompt section at order 100, exposes a profile tool with calibration metadata, and shows a realtime armor-status badge driven by a session projection.
liangmianya/dsh-synapse★ 463
Visual, non-linear conversation workspace for DeepSeek Harness — sessions, follow-ups and branches become a browsable conversation map.
ranxianglei/billion-context★ 450
The official billion-context plugin: a context-compression plugin for small context windows (a 100K context is enough), token savings (5x fewer tokens), and month-long single sessions (billions of tokens).
Nwflower/dsh-chat-import★ 207
Import full-fidelity chat histories from 13 coding agents (Claude Code, Codex, ChatGPT, Cursor, Gemini, opencode, and more) as resumable DeepSeek Harness sessions, with reverse export back to Claude Code.
Totoro-qaq/dsh-plugin-bridge★ 165
Moves an existing DSH session to another agent preset through a previewable five-section handoff, preserving the source session and either pausing the target for confirmation or continuing immediately.
Anionex/dsh-turn-rewind★ 125
Rewind conversation and workspace state, powered by a persistent Change Ledger.
Community comments
Comments are public GitHub Discussions. Loading them connects to GitHub and Giscus; a GitHub account is required to post.