dsh-eval-regression
Manifest valid★ 1该仓库暂未提供项目说明。
dsh-eval-regression
A small, deterministic regression-evaluation plugin for DeepSeek Harness.
It registers evaluate_golden_output, a model-callable tool that compares supplied candidate output against required and forbidden fragments. It does not call a model, persist data, or claim semantic correctness. Its job is repeatable pass/fail evidence, not vibes-based architecture in a trench coat.
Why
Agent changes routinely regress answers that appear superficially acceptable. A stable corpus of expected fragments gives a cheap, transparent signal for release smoke tests and replayed transcripts:
- required fragments catch omissions
- forbidden fragments catch known bad claims or unsafe fallbacks
- per-case reports make failures reviewable
- deterministic scoring is suitable for CI thresholds
Install as a DSH plugin
dsh plugin --profile <profile> add github:aryswisnu/dsh-eval-regression
The package is a DSH bundle. Its cordis.patch.yml registers the tool automatically after the profile's base tool runtime.
For local development:
git clone https://github.com/aryswisnu/dsh-eval-regression.git
cd dsh-eval-regression
npm install
npm run build
dsh plugin --profile <profile> add .
Run a version-controlled suite in CI
The plugin also ships a small CLI. It reads a JSON suite, prints an evaluation report to stdout, exits 0 when every case passes, exits 1 when any case fails, and exits 2 for invalid input or usage errors.
{
"suite": "release-smoke",
"cases": [
{
"id": "grounded-answer",
"actual": "The result is 42. Source: benchmark.csv",
"includes": ["42", "Source:"],
"excludes": ["I cannot verify"]
}
]
}
npx dsh-eval-regression suites/release-smoke.json
# or, from this repository:
npm run evaluate -- suites/release-smoke.json
The report includes total passed and failed cases, a 0..1 score, and case-level missing or forbidden fragments. This makes the evaluation corpus ordinary, reviewable source code and makes a failed expectation fail the CI job.
Tool example
{
"suite": "release-smoke",
"cases": [
{
"id": "grounded-answer",
"actual": "The result is 42. Source: benchmark.csv",
"includes": ["42", "Source:"],
"excludes": ["I cannot verify"]
}
]
}
The canonical result includes total passed and failed cases, a 0..1 score, and each case's missing or forbidden fragments.
Boundaries
This is intentionally a narrow deterministic evaluator. It does not replace model-quality review, factual grounding, tool execution checks, or snapshot replay. Use it as one gate in an evaluation harness, then add stronger signals where the product needs them.
Development
npm install
npm test
npm run typecheck
npm run build
MIT License.
Comments
Loading…
From the same category
by deepseek-ai
DeepSeek Harness: Everything is a Plugin.
★ 234.3k
MIT
TypeScript
Sep 23, 2026
by nexu-io
🎨 Best DeepSeek Harness Design Plugin. The open-source Claude Design alternative. 🖥️ Local-first desktop app. 🖼️ Your coding agent becomes the design engine: prototypes, landing pages, dashboards,
★ 97.8k
Apache-2.0
TypeScript
Sep 23, 2026
by Tencent
Open-source LLM knowledge platform: turn raw documents into a queryable RAG, an autonomous reasoning agent, and a self-maintaining Wiki.
★ 29.3k
↓ 838/wk
NOASSERTION
Go
Sep 23, 2026
dsh plugin --profile web add @wxg-prc-cpg/dsh-weknoraby awesome-dsh-plugin
A curated list of plugins for DeepSeek Harness (dsh) · DeepSeek Harness 插件精选列表
★ 16.7k
CC0-1.0
Python
Sep 23, 2026
by zhu1090093659
DeepSeek Harness (DSH) Web Plugin Aggregation Ecosystem · Everything is a plugin, distributed via the Creative Workshop
★ 8k
Apache-2.0
TypeScript
Sep 23, 2026
dsh plugin --profile web add dsh-webby yjh051108
dsh-routing-suite — injector + router-standard kit: install the runtime injector first, then the task-aware reasoning-mode router preset (measured P1-P23).
★ 7.2k
MIT
JavaScript
Sep 18, 2026
dsh plugin --profile web add @dsh-external/dsh-super-injector