Fine-grained LLM verification for DSH agents: pairwise rewards over score-token logprobs, Probabilistic Pivot Tournament best-of-N selection, and per-step progress tracking.
Install
# from npm (prebuilt)
dsh plugin --profile web add dsh-llm-as-a-verifier
# from GitHub (first run asks for allowBuilds approval — follow the hint, retry)
dsh plugin --profile web add github:TaurenMountain/dsh-llm-as-a-verifier
Any plugin you install runs third-party code with your own permissions — it can read your files, use your credentials, and reach the network, and tool approvals don’t sandbox it. GitHub-sourced plugins also run build scripts at install time — pnpm blocks those until you allow them, so an install can stop with ERR_PNPM_GIT_DEP_PREPARE_NOT_ALLOWED or ERR_PNPM_IGNORED_BUILDS; dsh prints the exact key to add under allowBuilds in your profile’s pnpm-workspace.yaml, and the install works on the next run. Allowing a build is a trust decision: only install sources you trust, and pin a commit (github:owner/repo#sha).
README
English | 中文
Brings LLM-as-a-Verifier — the unified verification framework — into DeepSeek Harness (
dsh) as a tool plugin. The agent can now verify its own candidates with fine-grained probabilistic feedback: instead of a binary good/bad judgement, the verifier's full logprob distribution over a 20-point score scale is read and its expectation taken.
What you get
Three model-facing tools, registered automatically after install:
| Tool | Capability | Cost |
|---|---|---|
verify_compare |
Score two candidates (code/plan/trajectory) against criteria, returning fine-grained rewards (scoreA, scoreB) in [0,1] |
1 verifier call per criterion per evaluation |
verify_select |
Best-of-N via the Probabilistic Pivot Tournament: O(Nk) comparisons instead of an O(N²) round-robin | Linear in N |
verify_track |
Per-step progress curve scored on the A(0%)..T(100%) scale at each checkpoint | O(K) calls regardless of trajectory length |
Why finer than LLM-as-a-Judge? The upstream framework's key ideas are: ① fine-grained scoring granularity (20-letter scale); ② expectation over the full logprob distribution of score tokens; ③ reliability via repeated evaluations and criteria decomposition. Building on its open-source implementation, this plugin adapts score extraction, pairwise prompts, the pivot tournament, and progress tracking, adds a logprob backend and token accounting, and packages them as a Cordis tool plugin for DSH.
Install
dsh plugin --profile web add dsh-llm-as-a-verifier
Requires dsh ≥ 0.1.0-rc.6 and Node ≥ 18. Restart dsh web (or wait for HMR).
Configure the verifier backend
The verifier model must be an OpenAI-compatible service returning token-level logprobs: DeepSeek's hosted API, a local vLLM/SGLang server, OpenAI, etc.
In your profile config (~/.dsh/profiles/<name>/cordis.patch.yml or ~/.dsh/cordis.patch.yml):
- id: llm-verifier
config:
baseUrl: https://api.deepseek.com # or vLLM: http://localhost:8000/v1
apiKey: '${DEEPSEEK_API_KEY}' # environment variables preferred
model: deepseek-v4-flash # omitted: DeepSeek defaults to deepseek-v4-flash, others auto-probe /models
maxConcurrency: 8
Credential resolution order (upstream parity): plugin config → OPENAI_BASE_URL + OPENAI_API_KEY → DEEPSEEK_API_KEY (implies the DeepSeek endpoint with thinking enabled). With no credentials configured, tool registration still works and calls fail with MissingAPIKeyError only when executed.
export DEEPSEEK_API_KEY=sk-... # the simplest setup
Usage
Once installed, just ask the agent in the conversation:
I wrote three candidate implementations. Use verify_select with
"correctness" and "performance" criteria to pick the best one, then use
verify_track to check whether my earlier fix steps made progress.
Configuration reference
| Config | Default | Description |
|---|---|---|
model |
DeepSeek: deepseek-v4-flash; else auto-probed |
Verifier model name |
baseUrl |
inferred from credentials | OpenAI-compatible endpoint |
apiKey |
inferred from environment | Prefer environment variables |
timeoutMs |
60000 |
Per-request timeout (ms) |
maxConcurrency |
8 |
Max in-flight verifier calls |
deepseek |
inferred from baseUrl | Force the DeepSeek call path (thinking enabled) |
prefill |
true |
Prefill the score tags on non-DeepSeek servers (more reliable letter distribution on vLLM/SGLang) |
compare / select / track |
true |
Register the corresponding tool |
Tool arguments (nEvaluations, pivots, seed, groundTruthNote, ...) mirror the upstream llm_verifier Python package; see the user guide and SOP.
Library use
import { Verifier } from 'dsh-llm-as-a-verifier'
const verifier = new Verifier({ baseUrl: 'http://localhost:8000/v1' })
const { scoreA, scoreB } = await verifier.compare(problem, a, b, { Correctness: '...' })
const result = await verifier.select(problem, candidates, { Correctness: '...' }, { pivots: 2 })
const curve = await verifier.track(problem, steps, { checkpoints: [1, 3] })
Development
npm ci
npm run check # typecheck + vitest (76 cases incl. end-to-end against a local mock logprobs server)
npm run build
License
This project is released under the MIT License.
Acknowledgments
Portions of the scoring and verification implementation are adapted from LLM-as-a-Verifier. Thanks to its authors and contributors; the relevant copyright and license notices are preserved in LICENSE.
Links
More in this category
Tencent/WeKnora#dsh-weknora★ 32192
Four read-only tools over a WeKnora knowledge base: list knowledge bases, hybrid passage search, reassemble one document's chunks in order, and WeKnora's own cited RAG or ReAct-agent answer with a resumable session id.
superdesigndev/treg★ 4192
Tool catalog for agents: search ~2,600 external endpoints (SEO and SERP, backlinks, social, people and company enrichment, ad libraries, scraping) by the task you want done, read each one's parameters and per-call price, then call it with the credential injected server-side. Ships the skill plus an MCP row that stays disabled until TREG_TOKEN is set.
TencentCloudBase/CloudBase-AI-Toolkit#dsh-plugin★ 1133
Tencent CloudBase backend for DeepSeek Harness — scaffold and deploy full-stack apps from chat, render query results as table cards with paging, sorting and CSV export, preview a deployment on its domain, and call the CloudBase MCP toolset (`mcp__cloudbase__*`) with device-code login.
gitroomhq/postiz-agent#dsh-postiz★ 503
Connects DeepSeek Harness to Postiz over MCP: list connected social media channels, fetch per-platform posting rules, and schedule, draft, or publish posts to X, LinkedIn, Instagram, Facebook, Threads, TikTok, YouTube, Reddit, Bluesky, Mastodon, Discord, Slack, Telegram and more; adds a postiz workflow skill.
EthanYoQ/Invoice-Downloader#dsh-invoice-downloader★ 491
Local IMAP invoice download, OCR, archive, and Excel reimbursement summaries for DeepSeek Harness.
anysearch-team/anysearch-dsh★ 447
AnySearch-powered real-time web and vertical search provider for DeepSeek Harness.
Community comments
Comments are public GitHub Discussions. Loading them connects to GitHub and Giscus; a GitHub account is required to post.