QA pipeline as 10 agent skills: requirement analysis, test strategy, case writing and review, E2E (Playwright) and API automation, exploratory, regression scope and bug analysis, backed by a shared knowledge base with format rules, risk model and schema validators.
Install
# from npm (prebuilt)
dsh plugin --profile web add dsh-qa-skills
# from GitHub (first run asks for allowBuilds approval — follow the hint, retry)
dsh plugin --profile web add github:fishzjp/qa-skills
Any plugin you install runs third-party code with your own permissions — it can read your files, use your credentials, and reach the network, and tool approvals don’t sandbox it. GitHub-sourced plugins also run build scripts at install time — pnpm blocks those until you allow them, so an install can stop with ERR_PNPM_GIT_DEP_PREPARE_NOT_ALLOWED or ERR_PNPM_IGNORED_BUILDS; dsh prints the exact key to add under allowBuilds in your profile’s pnpm-workspace.yaml, and the install works on the next run. Allowing a build is a trust decision: only install sources you trust, and pin a commit (github:owner/repo#sha).
README
English | 简体中文
Quick start
Install
Option 1: skills.sh cross-agent install (Claude Code / Cursor / Codex / OpenCode and 70+ other hosts, one command)
npx skills add fishzjp/qa-skills --skill '*'
However you install,
core/— the shared knowledge base dependency unit (not an executable skill) — must come along: installing any single skill without core breaks the relative-path references. To fix a partial install, re-runnpx skills add fishzjp/qa-skills --skill '*'(or copy thecore/directory manually). Option 1's--skill '*'full install is verified: all 12 skills + core land in place, references intact.
Option 2: the universal install script (auto-detects agent skills directories)
git clone https://github.com/fishzjp/qa-skills.git
cd qa-skills
./install.sh # interactive: auto-detects agent skills directories (~/.agents/skills, ...)
./install.sh --auto # or fully automatic
Option 3: the DeepSeek Harness (dsh) plugin (dsh-qa-skills on npm)
dsh plugin --profile web add dsh-qa-skills
- Manual install:
cp -r skills/* <your skills directory>/—core/must be copied along, every skill references it by relative path. - Verify:
ls <your skills directory>should show 12 skill directories +core/+qa-skills.VERSION. - Upgrade:
./install.sh --target <dir> --linkinstalls symlinks —git pullupdates in place. - Uninstall:
./uninstall.sh.
Skills are plain Markdown (frontmatter + relative-path references) with no host-specific dependencies:
| Host | Install directory | Status |
|---|---|---|
| Claude Code | ~/.claude/skills/ or <project>/.claude/skills/ |
✅ primary target; evaluations run on it |
| Shared directory | ~/.agents/skills/ |
✅ one copy for many agents (install.sh default) |
| DeepSeek Harness (dsh) | ~/.agents/skills/, ~/.dsh/skills/, or <project>/.agents/skills/ |
✅ verified end-to-end |
| Codex CLI | ~/.codex/skills/ |
🔶 not systematically evaluated |
| Other Skills-capable agents | their skills directory | 🔶 same |
The pipeline's per-stage context isolation relies on host sub-agent support; hosts without it degrade to sequential sessions joined by files — correctness is unaffected.
The qa-memory skill maintains a .qa/ knowledge base inside the project under test (Markdown entries + an index, committed with the project and shareable with the team), persisting environment quirks, flaky-noise judgments, defect patterns and other testing knowledge across sessions. Reading has two paths — the skill reads it when triggered; hosts with an entry line see it automatically every session:
| Host | Loading |
|---|---|
| Cursor / OpenCode / Codex / Gemini CLI / Windsurf / Devin, etc. | AGENTS.md entry line (native, zero config) |
| Claude Code | add @.qa/INDEX.md to CLAUDE.md (or via the @AGENTS.md import chain) |
| Aider | set read: .qa/INDEX.md in .aider.conf.yml, or pass --read |
| Other hosts | the skill loads it actively when triggered (fallback path) |
The entry line is written by qa-memory's bootstrap flow after your confirmation, or added manually. Every entry passes the memory_validate.py gate (schema / budgets / secret scan / poisoning defenses).
First run
Tell your agent:
Test this requirement: {description + repo URL}
The full pipeline runs from requirement understanding through risk and test-type decisions to the test report. For a single stage only (write cases / review / convert to automation / regression scope), just describe the need.
What it does
| You say | The framework does | Output |
|---|---|---|
| "Test this requirement" | qa orchestrates the 9-stage pipeline with human checkpoints |
Full QA assets + test report |
| "Write test cases from this PRD" | Code-first: requests the repo, reads the implementation, finds latent bugs, then writes | Dual-track cases: markmap (human) + schema.yaml (machine) |
| "How should we test this?" | Risk Map (evidence-backed ratings) → two-domain decisions: functional + 10 test types | 测试策略.md (incl. type_scope + handoff packages) |
| "Review these existing cases" | Independent review: testable-point denominator + coverage + executability | Revised case file + review record |
| "Convert cases to automation" | Page Object conventions, listeners-before-actions, assertion triple-check, zero-baseline-no-delivery, self-built data & cleanup | Runnable Playwright / pytest / k6 code |
| "Root-cause this bug" | Reproduce → read code to the line → 5-dimension impact analysis → regression advice | Bug entry (root cause / evidence / regression) |
Also usable standalone: exploratory-testing (charter-driven), api-testing, bug-analysis, regression-testing (diff → regression scope).
测试用例_markmap.md is plain Markdown (markmap syntax): the VS Code Markmap extension, npx markmap-cli, or markmap.js.org/repl render it.
This framework focuses on decision-making and execution for system-level black-box testing. The following are explicitly out of scope, each for a stated reason:
- Unit / integration testing — a development-side responsibility; the risk ratings and "covered" conclusions in test strategy assume existing safeguards at that layer (flagged in reports when unverified)
- Real-device mobile automation — the compatibility matrix currently covers desktop browsers; cloud device farms are a candidate for future expansion
- Frontend component testing / frontend performance automation — candidate directions, pending decision-layer validation
- Penetration testing / SAST & dependency scanning — pentesting requires professional hands-on expertise and authorized environments (moved to security specials); SAST is a dev-side CI tool, used only as a signal source for the business-security axis
- Chaos-engineering toolchains — fault injection ships as design method plus execution prerequisites; specialized toolchains are not bundled
Design
Executable cases
AI-written cases often look professional but cannot be executed — vague verdicts, placeholders, no time bounds, invented entry points. The single output standard: a person who has never read the requirements, with no walkthrough, can start working from the file alone. The same requirement, from this framework:
> Precondition: operator logged in, at 「营销中台 → 券工场 → 活动列表」
- **TC-03-05 Auto-close at end time** [P1]
- Steps: 1. Pick a published coupon ending in 10 minutes 2. wait for expiry
- Expected: status becomes 「已结束」 within 1 hour; past 1 hour = fail
Backed by 8 hard rules in skills/core/executability.md; a veto metric in evaluation — a non-executable case scores zero no matter its coverage.
Three-layer architecture: fewer instructions, stronger following
Stuffing methodology, templates, and rules into one SKILL.md reduces the rules an agent actually follows (per Red Hat's ACE practice notes, performance degrades beyond ~500 lines). The fix is a three-layer architecture:
L1 SKILL.md header Trigger boundaries: when to use, when not to, who hands off to whom
L2 SKILL.md body Workflow: the backbone every invocation walks (≤500-line ceiling)
L3 references/ + core/ Methods/rules/templates: loaded on demand, explicitly referenced
SKILL.md keeps only the workflow; everything else is pushed down and loaded on demand — the agent faces only the instructions it needs at each step.
Test-type decision matrix: deciding what not to test
Without the skill, models produced zero explicit type decisions across 30 evaluated samples (two model tiers) — prose that mentions performance and security but never decides which types to include, how deep, or what to exclude. Mentioning is not deciding.
The fix is the test-type decision matrix: ten test types, every axis must be answered — include requires signals, exclusion leaves an auditable trace, full depth has a budget cap; every decision lands in a machine-checkable type_scope. Measured: type recall on the weakest model 0 → 0.88 (see measured results).
How it works
Files are the pipeline state — every stage persists its output to disk; stages consume files, not session memory. Long pipelines don't depend on context; interrupted runs resume from files in a fresh session:
PRD / Code
│ requirement-analysis
▼
需求模型.md ·················· ⏸ clarification checkpoint
│ test-strategy (risk → two-domain decisions)
▼
测试策略.md (Risk Map + ten-axis type_scope) · ⏸ budget call
│ test-case-writing
▼
Cases: markmap (human) + schema.yaml (machine)
│ test-case-review
▼
⏸ execution-strategy call (manual / Playwright / API)
│ automated-e2e-testing / api-testing
▼
Execution artifacts + bug evidence → bug-analysis → regression-testing
▼
回归清单.md → 测试报告.md
- Evidence & risk models — every finding carries an evidence level (E0–E4); risk ratings without evidence are invalid, and the chain evidence → risk → strategy → cases is traceable end to end.
- Test-type decision matrix — ten axes, every one answered; include/exclude decisions leave an auditable trace; greppable signals are scanned into a prefill so weak models revise instead of generating from blank.
- Human-in-the-loop checkpoints — clarifications, execution strategy, bug triage, and budget calls are your decisions; the agent proposes, never decides. Once recorded, later stages cannot overturn them.
Measured results
Evaluated on 12 tasks: same model, same evaluation pipeline; the only difference is whether this framework is injected. Numbers come from the heterogeneous-judge re-evaluation and are reported as measured, including the adverse ones. Full methodology and raw data live in the locally maintained evaluation pipeline and are not distributed with this repo; milestone releases ship a cross-model gain-matrix snapshot (Releases), and the On/Off output comparison is in examples/:
| Metric | Without | With |
|---|---|---|
| Case-conformance score | 0.26 | 0.98 |
| E2E real execution (single task × 3 samples) | 0/3 runnable | 1 full + 2×(2/3) |
| Planted-bug detection | — | 75% |
| Quality (LLM judge) | 0.70 | 0.76 |
| API real-execution pass rate † | 100% | 99.2% |
| Token cost | 1× | 3.3× |
Test-type decision matrix, first round (2026-08-23, not yet in the formal gain table) — 5 test-type decision tasks (reference answers dual-annotated), weakest model deepseek-v4-flash (n=3): without the skill, zero explicit type decisions (0 even under lenient parsing — the blind spot is decision discipline, not type knowledge); with the skill, type recall 0 → 0.88, and code-signal-only axes absent from the PRD 0 → 8/9. Both numbers enter the formal table after task-pool growth and cross-model rounds.
- Case-conformance score: format × content-rubric composite, no judge; format-free samples score 0 (same caliber), the gap is driven primarily by format adoption; replicated across two generator models (0.20→0.99); the earlier 0.77 was pre-fix — errata in the CHANGELOG.
- E2E real execution: real browser + real app, no judge; the without-skill group mixes no-code and failing-code outcomes.
- Planted-bug detection: heterogeneous-judge caliber (100% under same-family judge).
- Quality: heterogeneous judge; Δ +6.1pp (95%CI includes zero; significant under same-family judging).
- API real-execution pass rate †: clean re-verification caliber (main model glm-5.2, n=3): 100% without vs 99.2% with — within the noise band, at parity; weak-model tier same direction (0.30 / 0.67, skill better). The earlier adverse result was traced failure-by-failure to evaluation-side defects, not skill defects — errata in the CHANGELOG.
- Token cost: better but more expensive — total-token ratio (per-task mean, skill fully injected): 3.3× on the main-model round, up to 9.5× on the weak-model round; a single-file ablation shows the gains cannot be obtained by taking just the core standards document.
Pre-registered gates: 4/7 under the same-family judge, 5/8 under the heterogeneous judge (different compositions incl. a sign flip). Coverage gains (heterogeneous judge): +8.7pp case-writing tasks (CI [0.5, 15.4]), +13.2pp all tasks (CI [2.8, 26.3]), +9.7pp defect detection (CI [3.3, 16.4]) — all significant; same-family figure +3.8pp (judge leniency quantified and corrected — see the CHANGELOG). An early +29pp single-sample estimate was shown to be noise.
Validity boundary: the with-skill evaluation mode pre-injects all skill instruction files (real hosts load on demand), so with-skill numbers are an upper bound — an in-situ probe (n=1) observed no decay; pairwise judging exceeded tie limits under all three judges (win rate voided — mechanism issue).
Documentation
- examples/ — Skill On/Off output comparison on the same PRD
- CHANGELOG.md — release history (milestone releases ship a gain-matrix snapshot)
- RELEASING.md (Chinese) — release rules and checklist
- Design & planning documents (DESIGN / decision-layer design / v2 blueprint) — maintainer-local, not distributed with this repo
skills/ the product (12 skills + shared core/)
qa/ orchestration entry (thin, no domain knowledge)
core/ shared knowledge base (installed as a dependency alongside skills, no task triggering): evidence / risk-model /
executability / testing-principles / report-template / case-format / coverage /
schema-extraction / clarify-pattern / test-type-matrix (decision matrix) /
triage (failure triage) / pipeline-integration (headless & CI conventions)
+ methods/ (5 design-method guides) + scripts/ (schema validator + type-signal scanner)
requirement-analysis/ test-strategy/ test-case-writing/ test-case-review/
automated-e2e-testing/ api-testing/ exploratory-testing/ bug-analysis/ regression-testing/
qa-memory/ test-reliability/ (flaky & suite-reliability governance)
.dsh/ dsh plugin trio (manifest in package.json's dsh.bundle)
assets/ visual assets (README hero images, share image og.jpg, social preview) + landing-page self-hosted fonts in fonts/
examples/ Skill On/Off output comparison
scripts/ gate scripts (validate_skills.py architecture red lines + validate_repo.py repo-level gate)
tests/ regression tests & installer smoke (test_product_scripts / test_memory_validator /
test_repo_gates / install_smoke.sh)
index.html website landing page (GitHub Pages build source)
Community
- Contributing guide — architecture red lines; local checks
python3 scripts/validate_skills.py+python3 scripts/validate_repo.py(same as CI) - 💬 Discussions for Q&A; Issues for confirmed bugs and concrete requests
- 🛡️ Security: private reporting per SECURITY.md
- 📜 Code of Conduct · 📋 CHANGELOG
License
Links
More in this category
zhu1090093659/dsh-web#packages/dsh-skill-explorer★ 8488
Skill center for the dsh web GUI: browse all loaded skills grouped by source, enable or disable model invocation, create new skills, and delete into a recoverable trash.
GanyuanRan/Aegis★ 1327
Software-engineering method pack for coding agents, with skills for baseline-first planning, systematic debugging, prompt hygiene, verification before completion, and repair/retirement tracking.
superdesigndev/superdesign-skill★ 630
Design skill for UI and marketing graphics on the Superdesign canvas: reads the repo for context, extracts its design system, then generates and iterates branchable design drafts, flow pages, and reusable components through the Superdesign CLI.
linhay/harmony-next.skills★ 361
HarmonyOS NEXT skill bundle for DeepSeek Harness with offline API references and DevEco, HDC, and emulator automation guidance.
dhicoc/dsh-reverse-skill★ 220
Complete reverse-skill pack (85 SKILL.md) as a DeepSeek Harness Cordis plugin: reverse engineering, authorized pentesting and security-research skill router.
sandbaseai/sandbase-skills★ 201
Mounts 88 packaged research, social-intelligence, marketing and business Agent Skills into dsh through the filesystem Skill provider.
Community comments
Comments are public GitHub Discussions. Loading them connects to GitHub and Giscus; a GitHub account is required to post.