dsh-verified-progress
Manifest validVerified-progress ledger for long-running DSH work: an independent verifier certifies each claimed step, and a fingerprint check voids any verdict whose verifier mutated the workspace.
dsh-verified-progress
A verified-progress ledger for long-running DSH work — where a verdict only counts if the thing that produced it provably stayed read-only.
A long-running task fails in a specific way: the agent reports progress it cannot prove, the report enters the transcript, and the next turn builds on a claim that was never checked. Nothing downstream can tell a verified step from an optimistic one, so the error compounds until the final answer is confidently wrong.
This plugin adds the missing check. When the agent claims a step is done, an independent verifier with a fresh context inspects the real workspace and returns a structured verdict. Only a clean, aligned, complete verdict enters the ledger. A rejected step stays in the ledger as evidence — it is never promoted to progress.
claim ──▶ independent verification ──▶ verdict
│
complete + clean + aligned ──▶ ledger (progress)
│
anything else ──▶ ledger (evidence only)
The part that is not a prompt
Read-only verifiers are usually enforced by asking the verifier to behave and by removing the write tools from its tool set. Both are prevention, and both are defeated by anything the tool filter does not name.
This plugin detects instead. The workspace is fingerprinted — every file's
size and mtime, plus a SHA-256 for files under the hash limit — immediately
before and after each verification episode. If anything changed, the verifier
mutated the workspace during a read-only audit, so its own findings are voided by
construction: the round is recorded as blocked / violation and can never be
promoted to progress, no matter how clean its report reads.
That matters because a verifier that writes is not a neutral witness. It can create the evidence it then reports, and no amount of prompt discipline or tool allow-listing proves it did not. A fingerprint comparison does not have to trust the verifier at all — which is why the ledger's admission rule can be absolute:
integrity === "violation" || contract !== "aligned" ⟹ status !== "complete"
This is enforced in code, not requested in a prompt. A verifier that omits a
control line degrades conservatively — an unstated integrity is read as
suspect, never as clean — because a placeholder must never be the thing that
certifies work as done.
What it does
1. Independent verification instead of self-report. The verifier is a separate subagent that inherits no conversation context. It reads the files and answers with three control lines:
Status: complete | incomplete | blocked
Integrity: clean | suspect | violation
Contract audit: aligned | unknown | needs_revision | invalid
2. A durable verified-progress ledger. Accepted steps are appended to
ledger.jsonl (one JSON object per line, fsync'd per append) with a last-wins
projection in state.json. A crashed process, a compacted context, or a fresh
session can reopen the run and answer: what is verified, what was rejected and
why, and what is left. Rejected rounds are preserved as evidence rather than
discarded, so a later turn can see which claims were already tried and failed.
3. Mutation detection. Described above; it is the admission rule for the ledger, not a side feature.
Tools
| Tool | Purpose |
|---|---|
verified_progress_verify | Verify one claimed step against the workspace; records the verdict in the ledger. |
verified_progress_ledger | Read back verified progress, rejected claims, and what remains. |
verified_progress_state | Current run state: verified rounds, pending claims, open items. |
Declaring what the guard may judge
verified_progress_verify takes a guard_scope: a workspace-relative path or
glob whose non-glob prefix becomes the fingerprint root.
verified_progress_verify claim="the parser handles CRLF input"
guard_scope="src/parser/**"
This is not optional decoration. The guard originally fingerprinted the whole
workspace root, and in the desktop profile that root is the harness home itself —
sessions/**, storages/** and plugin state are written continuously by the
very session the plugin runs inside. The first two verifications came back
workspace mutated (+3 ~5 -3), and every changed path belonged to the harness.
The verifier had written nothing. A guard that accuses the verifier of the
host's writes fails every round while looking like a conservative verdict, so
the scope is now declared rather than assumed:
- the scope must resolve inside the workspace; absolute paths and
..are rejected instead of silently widened; - when the workspace is the harness home, a scope is required, and
harness-owned roots (
sessions,storages,logs,telemetry,cache,profiles, …) are excluded only at the root, so a real project containing a nestedsessions/is still guarded; - the result prints
Fingerprinted: <path>, so a verdict never leaves its scope ambiguous.
Install
dsh plugin --profile web add dsh-verified-progress
Or install a released tarball, which needs no registry:
dsh plugin --profile web add \
https://github.com/lion231226/dsh-verified-progress/releases/download/v0.3.1/dsh-verified-progress.tgz
Then restart the harness: a plugin that failed to activate is not retried, and the module cache holds the previous version until the process restarts. No build step — the package ships plain ESM and has no runtime dependencies.
Storage
Everything lives under the harness state directory, not in your project:
<state>/verified-progress/runs/<runId>/ledger.jsonl
<state>/verified-progress/runs/<runId>/state.json
<runId> is derived from the session id when there is one, otherwise from a hash
of the working directory, so reopening the same session reopens the same run.
Design limits — read these
- The verifier is a separate agent, not a separate process. It cannot see the executor's conversation, which is what makes it independent, and it runs with a read-only tool set. It runs inside the harness and is subject to the same permissions; the fingerprint check is what makes a violation detectable rather than impossible.
- The mutation guard is a correctness tool, not a security boundary. It compares file fingerprints. It does not defend against a local attacker, it is not a sandbox, and it does not cover paths excluded from the snapshot.
- Unhashed large files are reported as an evidence gap. Files above the hash limit are compared by size and mtime, and the snapshot says so rather than claiming a comparison it did not make.
- Verification costs tokens. Each verified step adds one verifier episode. The ledger is what you get for it; if your task is a single short turn, you do not need this plugin.
Why this exists
The loop design is ported from LongHorizon-Harness (AMAP-ML, MIT) — manager/executor/auditor rounds, a durable round ledger, and an auditor that cannot certify a dirty result.
That project is a Python orchestrator that wraps an agent CLI. It does not run on
Windows: its persistent layer requires os.O_NOFOLLOW, os.O_DIRECTORY and
supports_dir_fd (all absent on win32) and it shells out with POSIX
VAR=value cmd templates. It also, by its own design, drives one agent CLI per
role episode — so on DeepSeek Harness it can only read the final answer of each
dsh --profile headless run, not the intermediate tool events.
This plugin takes the parts that carry the value — independent verification, the downgrade invariant, the verified-state ledger, the mutation guard — and implements them natively in the harness, where the events are available and no POSIX-only primitive is needed. Its workspace-mutation detection is its own contribution; the upstream project guards against verifier writes only through the same prompt-and-allow-list prevention described above.
Verifying this plugin
npm test # 37 tests: verdict grammar, guard, ledger, tools, extraction
npm run verify # acceptance probe against the packed artifact
npm run verify is not a unit test run. It rebuilds the package from the
files list in package.json, loads the tool set out of that artifact, and
runs three scenarios against a real file system with a scripted verifier:
| Scenario | What it must show |
|---|---|
clean-verdict-is-progress | a clean verdict becomes progress, in one verifier episode |
positive-control-mutation-voids-verdict | a verifier that edits a file is caught; the round is blocked / violation and never becomes progress |
ledger-survives-a-fresh-context | a brand-new context reads verified progress back off disk |
The middle scenario is a positive control: if the fingerprint guard ever
stops noticing a mutation, positiveControlTripped turns false and the probe
exits non-zero, so a broken detector cannot be reported as a pass. The probe also
proves it ran — it records the artifact it loaded, the tool names it found, and
that all three scenarios reached a deciding assertion.
The result of the last run is committed: verify/acceptance-report.json
and the raw tool results and ledger contents in
verify/acceptance-evidence.json.
Requirements and known environment issues
-
A subagent provider that does not inherit the parent context. The plugin injects
toolsandsubagents; thespawnprovider shipped with the web and headless profiles satisfies this. If only theforkprovider is available, the plugin refuses to certify anything and says so in the tool result rather than silently falling back to a verifier that can see the executor's reasoning. -
Windows: a PowerShell that can actually be spawned. DSH resolves
pwshthrough a fixed candidate list (@deepseek-ai/dsh-pwsh-local); a Microsoft Store PowerShell can leave a zero-byte execution alias in%LOCALAPPDATA%\Microsoft\WindowsApps\pwsh.exe. That alias is a regular file tolstat, so it is selected, and Node then fails withspawn ...\WindowsApps\pwsh.exe ENOENT. If your harness will not start for that reason, point the resolver at a real binary:# profile cordis.patch.yml - id: pwsh-local config: pwshPath: 'C:\Program Files\WindowsApps\Microsoft.PowerShell_7.6.6.0_x64__8wekyb3d8bbwe\pwsh.exe'This is a host resolution issue, not a plugin one, but it is recorded here because it is the first thing that stops a Windows acceptance run.
License
MIT
Comments
Loading…
Similar plugins
by 863683348
Verification toolkit for DSH agents: evidence-based claim checking against workspace files with line citations, config validation (JSON/YAML), and read-only URL/npm/GitHub submission-readiness probes.
★ 0
MIT
JavaScript
Sep 11, 2026
dsh plugin --profile web add dsh-plugin-verifyby MaxHou-infinity
Evidence-driven company and job due-diligence for DSH: logs sources with evidence levels, adds evidence-bounded claims, verifies the legal entity against official registries, and renders a PROCEED/VER
★ 3
↓ 287/wk
MIT
TypeScript
Sep 4, 2026
dsh plugin --profile web add dsh-scoutby 82c86b8z86-stack
Engineering workflow layer for DeepSeek Harness (dsh): a disciplined-engineer agent preset with five gated phases — requirements clarification, plan approval, TDD, parallel subagent execution, and ver
★ 4
MIT
JavaScript
Aug 17, 2026
dsh plugin --profile web add dsh-engineering-workflowby TaurenMountain
LLM-as-a-Verifier for DeepSeek Harness: fine-grained reward, Probabilistic Pivot Tournament best-of-N selection, and per-step progress tracking as agent tools.
★ 5
↓ 106/wk
TypeScript
Aug 24, 2026
dsh plugin --profile web add dsh-llm-as-a-verifierby PerryLink
Verifiable research-report engine for DeepSeek Harness: content-addressed evidence ledger (claim-snapshot binding, tamper-evident) plus versioned sealed reports with per-claim verification verdicts an
★ 214
↓ 1.5k/wk
Apache-2.0
TypeScript
Oct 10, 2026
dsh plugin --profile agent add dsh-research-reportby ardesp0630
工作进度面板:在输入框下方的带区常驻显示当前会话的任务清单完成度与运行状态,展开可见任务明细、轮次步骤、模型/工具耗时与上下文占用。
★ 0
MIT
JavaScript
Sep 26, 2026
dsh plugin --profile web add dsh-work-progress