LongCat-2.0 模型适配器:100 万上下文、二元思考模式、工具调用,基于 OpenAI 兼容接口的流式输出。
安装
# GitHub 源码(首次需按提示配置 allowBuilds 构建授权后重试)
dsh plugin --profile web add github:ffyuuu/dsh-llm-longcat
装任何插件都等于在你的机器上跑第三方代码,权限和你本人一样大——能读你的文件、用你的凭据、访问网络,工具审批管不到它。GitHub 来源的插件还会在安装时执行构建脚本——pnpm 默认拦截,所以安装可能停在 ERR_PNPM_GIT_DEP_PREPARE_NOT_ALLOWED 或 ERR_PNPM_IGNORED_BUILDS;dsh 会打印出需要添加的确切键名,把它加进该 profile 的 pnpm-workspace.yaml 的 allowBuilds 下,重跑一次即可装上。放行构建本身就是一次信任判断:请只安装可信来源,并尽量锁定 commit(github:owner/repo#sha)。
README
该插件的 README 只有英文版本。
LongCat adapter for the DeepSeek Harness LLM seam.
Adds LongCat-2.0 as a model provider: 1M context, thinking mode, tool calling.
Features
- Thinking mode — recognizes LongCat's
reasoning_contentfield and translates it into harnessReasoningBlocks - Tool calling — full function-calling support, with
argumentskept a raw JSON string end to end - Multi-turn — replays
reasoning_contenton tool-call turns, as thinking-mode passback requires - Streaming — SSE with the
usage-before-finishordering the harness relies on - Credential seam — the key resolves per request from
ctx.credentialsor the environment; no secret in any config file
Supported models
| Model | Context | Max output | Notes |
|---|---|---|---|
LongCat-2.0 |
1,048,576 | 131,072 | text-only; thinking + tool calling |
Facts from GET /openai/v1/models/LongCat-2.0, the only documented endpoint that
reports supported_parameters. Tool calling is not mentioned on the
chat-completions doc page and is only visible there.
Install
dsh plugin --profile default add github:ffyuuu/dsh-llm-longcat
export LONGCAT_API_KEY=... # create one at https://longcat.chat/platform/api_keys
Installing a bundle lets the package's install scripts run on your machine, outside the sandbox the agent runs under. Pin a commit so a later push cannot change what executes:
dsh plugin --profile default add github:ffyuuu/dsh-llm-longcat#3dcb3b1b5870ba52baab053453bdbb28826e5f13
Then pick LongCat-2.0 in the model selector. The key may also be stored through the Web UI's Models page instead of the environment.
If dsh itself will not install
At the time of writing, installing the harness can fail before any plugin is
reached, with either ETARGET … dsh-typert-protocol@^0.1.0-rc.8 or an npm
heap exhaustion. That is an upstream packaging state, not this plugin:
@deepseek-ai/dsh published 0.1.0-rc.8 while several packages it depends on
stopped at 0.1.0-rc.7, and because the manifests use caret ranges,
^0.1.0-rc.7 still resolves up into the missing rc.8. npm then backtracks
over an unsatisfiable graph until it runs out of memory.
Pinning every @deepseek-ai/* package to an exact 0.1.0-rc.7 through npm
overrides avoids the drift. Nothing in this plugin needs changing either
way — it declares >=0.1.0-rc.7 and works against whichever of those the
host ends up with.
Config
- id: llm-longcat
name: dsh-llm-longcat
config:
apiKeyEnv: LONGCAT_API_KEY # default; resolved per request, never a literal key
baseURL: https://api.longcat.chat/openai/v1 # optional; $LONGCAT_BASE_URL then the public API
thinking: enabled # optional deployment policy; `disabled` locks every request to off
reasoningEffort: high # optional; off | high — LongCat's switch is binary
maxTokens: 131072 # optional per-request output cap
defaultContextWindow: 1048576
streamIdleTimeoutMs: 300000 # optional; five-minute default
retryPolicy: # optional; omission uses bounded normal defaults
mode: normal
maxRetries: 3
models:
- id: LongCat-2.0
contextWindow: 1048576
A llm-longcat: section in $DSH_HOME/settings.yaml overrides any field
without a restart: base URL, catalog, request defaults, and idle budget all
take effect on the next request, while an in-flight stream keeps the facts it
started with.
Reasoning is binary, deliberately
LongCat controls thinking with thinking: {type: enabled|disabled} and does
not accept OpenAI's top-level reasoning_effort — its
supported_parameters lists the former and omits the latter. There is
therefore no low/medium/high gradient to map, and this adapter offers exactly
two levels rather than advertising controls that would collapse onto the same
two request bodies:
| Selected effort | Wire body |
|---|---|
high ("Thinking") |
{"thinking": {"type": "enabled"}} |
off |
{"thinking": {"type": "disabled"}} |
| (none named) | resolves from config; still explicit |
off serializes an explicit disabled rather than omitting the field —
omitting it would hand the decision to LongCat's server-side default, which is
not what selecting Off should mean. Requesting low, medium, or max fails
with UNSUPPORTED_REASONING_EFFORT before any network I/O.
Wire-format notes
- Tool-call deltas repeat
idandnameas explicitnull. LongCat sends them on the opening delta and thennull(not omitted) on every continuation, so a naive!== undefinedguard blanks the assembled call's name. Verified on live traffic; pinned by a regression test. - Streaming only, with
stream_options.include_usagealways on. Usage may arrive attached to the finish chunk or as a trailing usage-only chunk; both are deferred to[DONE]sousagealways precedesfinish. - The first thinking-mode delta can be an empty string — it must not open a reasoning block.
- Reasoning passback: on assistant turns that carried tool calls,
reasoning_contentis serialized back into history; on tool-call-free turns it is dropped (ignored anyway — saves tokens). - Assistant
contentis always a string, never null: the message is durable session history, and a null there would make later turns replay a body the endpoint can reject. - Cache accounting:
prompt_tokens_details.cached_tokensmaps tocacheReadTokensand is subtracted out ofinputTokensto keep the harness's disjoint-count convention.
Errors
Non-2xx responses throw LlmError with stable codes. LongCat documents a
dedicated 402 for exhausted token quota and puts insufficient_quota on
403, where most OpenAI-compatible providers use 429 — both are classified
as QUOTA before the auth and rate-limit buckets, so a depleted balance is
never reported as a bad key or retried as a transient rate limit.
| Condition | Code |
|---|---|
| 402, or quota detail at any status | QUOTA_EXCEEDED |
| 401 / 403 | AUTH |
| 429 | RATE_LIMIT |
| 400 with context-overflow detail | CONTEXT_WINDOW_EXCEEDED |
| other 400 | INVALID_REQUEST |
| 5xx | SERVER |
no [DONE] / bad JSON |
STREAM_CLOSED / MALFORMED_RESPONSE |
A completed stream that opened no content blocks becomes a finish error with
EMPTY_RESPONSE, which the shipped retry policy treats as retryable.
Tests
npm run typecheck # against the published @deepseek-ai/dsh-llm types
npm test # 30 unit tests over serialize + translate
npm run build # emits lib/ and lib/types/
npm run test:e2e # real API, needs LONGCAT_API_KEY, spends a few hundred tokens
test:e2e drives the built adapter's own serialize → SSE → translate pipeline
against api.longcat.chat, so it verifies what the plugin actually sends
rather than a hand-written approximation. It is what caught the null-name
delta bug.
Limitations
- No image input. LongCat-2.0 reports
modality: text->text, so image content is refused before sending, naming the model. - No stop sequences.
stopis absent fromsupported_parameters; passing one fails withUNSUPPORTED_OPTIONrather than silently running past it. - Reasoning is binary — no low/medium/high gradient exists to map.
License
MIT
链接
同类插件
V1ki/dsh-plugin-subscriptions★ 421
把 ChatGPT(Codex)、Claude、Grok 订阅当作 DeepSeek Harness 的 LLM 提供方:设置页登录、模型目录、用量展示,以及 image_generate、video_generate 与 x_search 工具。
Mars-Sea/dsh-commandcode-provider★ 363
非官方 Command Code 模型接入插件:注册 `commandcode` 路由,带实时模型目录与推理强度支持。
corrinehu/dsh-workbuddy-connect★ 293
将 WorkBuddy 桌面 App 包含的模型自动接入 DeepSeek Harness,在 DSH 对话窗口里零配置使用。
cv-superding/dsh-deepseek-web-login★ 223
新增 deepseek-web provider,把 chat.deepseek.com 网页端模型接入 DSH:浏览器登录抓取、PoW 请求签名、SSE 流式传输与基于提示词的工具调用。
volcengine/ark-cli#ark-plan-api★ 139
在 DSH 原生模型选择器中注册方舟 Agent Plan、Coding Plan 与后付费模型路由。
franksong2702/dsh-codex-connect★ 135
通过 ChatGPT OAuth 将 OpenAI Codex 模型接入 DeepSeek Harness,并提供可选的搜索与图片工具。
社区评论
评论公开保存在 GitHub Discussions。加载评论会连接 GitHub 和 Giscus;发表内容需要 GitHub 账号。