按提供方/模型限制并发生成,并附带设置卡片:限定同时进行流式生成的会话数,在流式调用与子代理派生两处强制执行;超出上限的扇出会进入 FIFO 队列等待,而不是直接失败。
安装
# npm 包(预构建)
dsh plugin --profile web add @creait/dsh-gen-limit
# GitHub 源码(首次需按提示配置 allowBuilds 构建授权后重试)
dsh plugin --profile web add github:CREAIT-nl/dsh-plugins#path:/gen-limit
装任何插件都等于在你的机器上跑第三方代码,权限和你本人一样大——能读你的文件、用你的凭据、访问网络,工具审批管不到它。GitHub 来源的插件还会在安装时执行构建脚本——pnpm 默认拦截,所以安装可能停在 ERR_PNPM_GIT_DEP_PREPARE_NOT_ALLOWED 或 ERR_PNPM_IGNORED_BUILDS;dsh 会打印出需要添加的确切键名,把它加进该 profile 的 pnpm-workspace.yaml 的 allowBuilds 下,重跑一次即可装上。放行构建本身就是一次信任判断:请只安装可信来源,并尽量锁定 commit(github:owner/repo#sha)。
README
该插件的 README 只有英文版本。
Per provider/model concurrency limits for DeepSeek Harness, with a settings card.
Some backends fall over — or bill hard — when several sessions generate against them at once. A self-hosted GPU serving one model has a real ceiling; a metered API has a financial one. dsh has no per-model concurrency control, so a single agent that fans out subagents can saturate either.
This caps how many sessions may generate concurrently on a given provider/model. Work past the cap waits in a FIFO queue rather than failing: a fan-out of eight researchers against a limit of three is a pacing problem, and bouncing five of them does not make the brief smaller, it just spends the retry budget re-asking.
How it enforces
The count — llm/stream waterfall. Every streaming model call is capped by
the number of distinct sessions actually generating. A session reentering
the loop is not counted twice, and a session parked waiting on a subagent holds
no stream — the child is what counts. This is where the limit is enforced, and
it is enforced the same way no matter how a session started: a subagent from a
tool call, a child the workflow engine started directly, or a plain
conversation.
The pacing — tools/pre-execute. When an agent calls a subagent spawn tool
(subagent, subagent_fork) and the target provider/model is already full, the
spawn joins the same queue and is admitted as soon as there is room, so a
fan-out is paced instead of piling more sessions onto a saturated backend.
This gate deliberately holds no slot of its own. Reserving one per admitted
spawn is the obvious design and it double-counts: the child then takes a second
slot the instant it generates, so every live child costs two. There is no clean
seam to hand a reservation over either — subagent/start carries no
back-reference to the spawn that caused it, and it also fires for children the
workflow engine starts without any tool call. So a child is counted exactly
once, where it can be counted consistently.
What a wait costs. queueTimeoutMs bounds how long a request waits (0
waits indefinitely) and maxQueued bounds how many may wait at once — an
unbounded queue in front of a slow backend is a memory leak that presents as a
hang. Only a request that exhausts its wait fails, and it fails with the code
GEN_CAPACITY_EXCEEDED; for a spawn that means the tool call is denied.
Reaching that point means the backend has been saturated for a sustained
period, not that a request was unlucky with timing.
The consequence — transport timeouts. Waiting for a slot means a stream may
legitimately go quiet for a long time, so the limiter also makes sure the socket
agrees. llm-pi-ai lets a provider declare streamIdleTimeoutMs, but the SSE
stream rides Node's built-in fetch, whose bodyTimeout defaults to five
minutes and which nothing in the harness configures — so any value above
300000ms is unreachable. Raising it just moves the kill from the harness
watchdog (TIMEOUT) to undici (TypeError: terminated, classified TRANSPORT,
equally retryable), and each retry restarts the step from scratch and discards
everything it had generated.
So transport.js reads the timeout the provider already declares and
installs a dispatcher that applies it to that provider's origin, plus a
30-second margin so the harness watchdog stays the one that reports a dead
stream. There is nothing new to configure, and no other origin is affected — MCP
servers, web fetches and the update check keep Node's defaults. A provider that
declares no streamIdleTimeoutMs, or one under five minutes, is left alone.
Install
dsh plugin --profile web add @creait/dsh-gen-limit
The package ships its own cordis.patch.yml, so it inserts its roster row on
its own — no manual profile edit. Add it to dsh.profile.bundles to activate
the browser half.
Configure
Limits live in the dsh-gen-limit settings namespace, one row per
provider/model. max: -1 means unlimited, and any pair without a row defaults
to unlimited — the plugin is inert until you give it a limit.
Seed them from the row in your profile patch — provider and model are
whatever ids your own routes publish:
- id: gen-limit
config:
limits:
- { provider: local-gpu, model: deepseek-v4-flash, max: 2 }
- { provider: anthropic, model: claude-opus-4, max: 1 }
queueTimeoutMs: 120000 # how long a request waits for a slot; 0 = forever
maxQueued: 64 # how many may wait at once
Or edit it in the GUI: Settings → Plugins → Plugin config → Generation
Concurrency (the shipped web UI labels those 设置面板 → 插件 → 插件配置).
The card lists the live providers and models from the same llm service the
conversation uses, so the rows are pickable rather than typed from memory.
Routes
The card talks to three plugin-owned loopback routes rather than the settings RPC — the harness settings wire only exposes namespaces on its own allowlist, which a plugin cannot widen:
| Route | Purpose |
|---|---|
/api/dsh-gen-limit/config |
read/write the limit rows |
/api/dsh-gen-limit/catalog |
live provider/model list |
/api/dsh-gen-limit/stats |
what is generating right now |
What breaks this
llm/stream and tools/pre-execute are pre-1.0 internal seams with no
compatibility guarantee. peerDependencies pins the versions this was
built against; a harness upgrade can move them.
The transport half rests on the seam Node leaves for proxies: built-in fetch
takes no per-call timeout options and reads its dispatcher from a global that
undici's setGlobalDispatcher writes. Verified on Node 25.8.1 with undici
8.10.0 — a 3s bodyTimeout installed this way killed a built-in fetch body at
3.5s, where the default had taken 301s. It is a convention, not a contract; a
runtime that stopped honouring it would put the five-minute ceiling back, which
is where things stood before this existed.
The settings nav glyph is a deliberate reach past the API. settings.section
has no icon option — the shell picks the glyph from a hardcoded section-id map
(ui-settings-general navIcon) and falls back to the gear for ids it does not
know, ours included. So the client half repaints its own row: it finds the nav
cell by label and swaps the gear's path geometry for the official
IconBranchOutline16 path, mutating the attribute rather than replacing the
node so React re-renders over it without restoring the gear. It fails safe — if
the shell's markup moves, nothing matches and the row keeps the gear.
链接
同类插件
V1ki/dsh-plugin-subscriptions★ 426
把 ChatGPT(Codex)、Claude、Grok 订阅当作 DeepSeek Harness 的 LLM 提供方:设置页登录、模型目录、用量展示,以及 image_generate、video_generate 与 x_search 工具。
Mars-Sea/dsh-commandcode-provider★ 370
非官方 Command Code 模型接入插件:注册 `commandcode` 路由,带实时模型目录与推理强度支持。
corrinehu/dsh-workbuddy-connect★ 314
将 WorkBuddy 桌面 App 包含的模型自动接入 DeepSeek Harness,在 DSH 对话窗口里零配置使用。
cv-superding/dsh-deepseek-web-login★ 235
新增 deepseek-web provider,把 chat.deepseek.com 网页端模型接入 DSH:浏览器登录抓取、PoW 请求签名、SSE 流式传输与基于提示词的工具调用。
FishBottle7/opencode2dsh★ 150
将 OpenCode Zen 免费模型接入 DeepSeek Harness,无需 API Key。
WSL043/dsh-codex-subscription★ 145
通过 ChatGPT OAuth 在 DSH 中使用 Codex 模型,提供订阅联网搜索、额度与安全重置、图片工具、高速模式和模型感知上下文;无需 API Key 或 Codex CLI。
社区评论
评论公开保存在 GitHub Discussions。加载评论会连接 GitHub 和 Giscus;发表内容需要 GitHub 账号。