dsh-goalloop-agentteams
Manifest valid★ 1Deterministic goal-completion gate for DSH: contract-first AC grammar, re-runs every verify command itself (zero model trust), binds verdicts to a tree digest, denies false completion on update_goal/u
dsh-goalloop-agentteams
Deterministic goal-completion gate for DeepSeek Harness — ports the judging face of the goal-loop skill onto native Agent Teams + the task board, and drives the full agentic loop from one trigger.
Contract-first acceptance criteria, a gate that re-runs every check itself (zero model trust), verdicts bound to a tree digest, and a hard deny on false completion at update_goal / agent_teams_update_task / team_task_update via tools/pre-execute. The portable team-file protocol and dual-ledger bookkeeping from goal-loop are deliberately not ported — native Agent Teams is strictly stronger.
Trigger the loop: /goal-loop-at <objective> (or the goal_loop_at tool) writes the contract + round config, returns the loop protocol and suggested Agent Teams tasks derived from each AC. Iterate via goal_gate_check (score, trend, failedActions), claim completion only on GO.
中文说明见 README.zh.md。
Why
DSH already has two layers of completion judgement — but neither is a deterministic gate:
| Layer | Judges | Gap |
|---|---|---|
dsh-agent-teams quality kinds | Did each step follow its contract | verify commands are not re-executed by the runtime; commandsRun exit codes are member-supplied |
dsh-task-board acceptance gate | Does it look right semantically | LLM A/B comparison (threshold 0.65) — and every card can opt out via skipVerification |
dsh-goalloop-agentteams is the hard floor underneath both: it re-runs every AC's check command itself, binds verdicts to a tree digest, and counts caught false-completes — two catches → BLOCKED. Deterministic gate = floor, LLM judge = semantic layer, quality kinds = step contract. Stacked, they close the loop.
Contract grammar
objective: <one-line goal>
AC-1 | <yes/no statement> | check: `<command>` | [probe: `<probe>`] | [metric: `<regex with one capture group>`] | [baseline: delta|abs] | expected: <spec>
AC-2 | ... | check: `...` | expected: exit=0
Spec = exit=0 | <op><number> (<=5, >0, =3) | maximize | judged.
Every check must be a named, failable command — include the environment dependency, the empty case, the error path — or the gate degrades into a tautology.
[probe:]is verify-the-verifier (R9): the probe runs first; a probe failure marks the ACunverifiablerather thanpassed. More than 1/3 unverifiable → the whole gate returns NO-GO.[metric:]extracts the metric value from stdout (one capture group). Without it the gate falls back to the last number in the output — writemetric:whenever the output contains other numbers.maximizecompares against the previous round's metric:baseline: deltarequires strict improvement,baseline: abs(default) requires no regression. The first run has no baseline →unverifiable(rungoal_gate_checkonce to establish it).judgeduses theprobeas a deterministic judge — the probe's exit code is the verdict. Without a probe,judgedisunverifiable(Goodhart guard).- Contract lives at
.goal-gate/goal.mdin the workspace root.
Gate & interception
goal_gate_check re-runs every check itself:
rc=0GO /rc=2NO-GO /rc=3BLOCKED /rc=4state error- Output carries the optimization signal:
score(passed/total),round/maxRounds/remainingRounds,bestScore,regression(score dropped below the high-water mark),trend(last 5 rounds),failedActions(per-AC repair list) - Interception is mounted on
tools/pre-execute, coveringupdate_goal(action:'complete'),agent_teams_update_task(status:'completed')andteam_task_update(action:'complete')(legacyupdate_task(status:'completed')kept as alias)
Goal loop (right loop / right eval / right metric)
/goal-loop-at <objective> (slash command) or goal_loop_at (tool) starts the loop:
- writes the contract skeleton (placeholder checks are fail-closed
TODO-REPLACE-ME) and.goal-gate/loop.json(maxRounds, default 8); - returns the loop protocol and suggested Agent Teams tasks derived from each AC;
- every gate evaluation is appended to
.goal-gate/history.jsonl(ts/trigger/round/code/score/totals/failed ACs) — the trajectory the loop optimizes against; - NO-GO → turn
failedActionsinto repair tasks, re-check; never claim completion before GO. Two caught false-completes →BLOCKED(human). Round budget exhausted →roundsExhausted, stop and escalate.
Self-optimization rule: score must not regress (regression: true → fix the regression first); maximize ACs with baseline: delta require strict metric improvement round over round.
Three hard constraints (measured, not assumed)
From an ordering probe (tools/pre-execute waterfall semantics, real deny-return shape):
- Waterfall is first-registered-first-block. Only the earliest registered layer's deny reason reaches the user; later layers are never called. So the reason must carry its own
code([goal-gate:no-go]/verdict-stale/blocked) — otherwise the user can't tell which layer blocked. denyandthroware different propagation mechanisms and must not be mixed.dsh-agent-teams's quality-gate rejects bythrowing, which blows through the waterfall and hands the caller an exception; the task board returns a structured{kind:'deny'}. This plugin always uses structureddeny— neverthrow— so its error path stays distinct from quality-gate exceptions.- Digest binding must be self-built. DSH has no verdict↔file-tree binding anywhere (only state-key hashes). This plugin stores the digest and re-computes it before denying, so a passing verdict is voided automatically when the tree moves (R7).
R7 digest caveat (fixed)
If treeDigest walks the ledger directory, writing the ledger changes the digest → the gate spuriously invalidates its own verdicts and R1's false-complete counter never reaches 2. Fixed: the digest excludes .goal-gate, node_modules, .git.
Interpreter constraint
The DSH Host bundles Node 24.21.0 (runtime/primary-runtime/dependencies/node/bin/node). A node on PATH may be broken (e.g. SIGKILL on launch). Plugins run inside the Host runtime — never write a bare node in a gate script; if you must spawn a child, resolve it via config.node ?? process.execPath (the pattern dsh-skill-office uses).
Install
dsh plugin add dsh-goalloop-agentteams
Or from source: the repo declares dsh.bundle in package.json with a cordis.patch.yml beside it, so dsh plugin add picks it up directly. apply(ctx, config) registers the goal_loop_at / goal_gate_init / goal_gate_check tools, the /goal-loop-at / /goal-gate commands, and the tools/pre-execute listener.
Tests
node test/run-local.mjs
The real @deepseek-ai/dsh-tools lives inside the DSH Host and is not importable outside it, so the runner injects a minimal defineTool shim from test/fixtures/dsh-tools-shim for the duration of the run.
Mapping to goal-loop
| goal-loop | This plugin |
|---|---|
| AC contract grammar + sha1 stamp | parseContract / contractStamp (R3) |
goal_gate.sh --check re-run | runGate (R9 probe + expected-spec judgement) |
| digest-bound verdicts (R7) | treeDigest + verdict-stale deny |
| R1 false-complete count → BLOCKED | falseCompleteRule |
/goal-loop-at trigger + loop orchestration | goal_loop_at tool / /goal-loop-at command + loop.json round budget |
| metric trajectory / self-optimization | history.jsonl + score / bestScore / regression |
goal_team.sh portable layer + dual ledger | dropped — use native Agent Teams + task board |
License
MIT
Comments
Loading…
Similar plugins
by zriyox
一句话大需求先别急着跑:拦一下问你要不要先规划,自动生成带验收清单的 plan.md 再开工 — DeepSeek Harness 插件 / Catches one-sentence mammoth tasks and scaffolds plan.md + a capped goal first
★ 3
MIT
TypeScript
Aug 15, 2026
dsh plugin --profile web add dsh-goal-scaffoldby leeyoung1
Pairs the running executor with a stronger reviewer model through zero-parameter advisor() calls and step-based patrol checks, returning plan, correction, or stop guidance.
★ 3
↓ 318/wk
MIT
TypeScript
Sep 11, 2026
dsh plugin --profile web add dsh-advisor-pluginby gosomea
DSH task supervision with DAG progress, plan review, persistent control and stage/completion checks.
★ 0
NOASSERTION
TypeScript
Oct 5, 2026
dsh plugin --profile web add dsh-task-supervisorby Cavan-Ou
Battle-tested multi-agent collaboration playbook for DeepSeek Harness: model-tier routing, spec discipline, git single-writer rule — as an installable skill. 多 agent 管线里运行 DSH 的实战协作规范
★ 7
MIT
JavaScript
Sep 11, 2026
dsh plugin --profile web add hermes-dsh-collabby dfycaly98931680
Agent trajectory governance and anomaly diagnosis: rebuilds flat session logs into multi-branch trajectory trees, detects loop deadlock, invalid retry, and goal drift, alerts with cost attribution, on
★ 3
MIT
TypeScript
Aug 16, 2026
dsh plugin --profile web add dsh-trajectory-governanceby lion231226
Verified-progress ledger for long-running DSH work: an independent verifier certifies each claimed step, and a fingerprint check voids any verdict whose verifier mutated the workspace.
★ 0
MIT
JavaScript
Oct 9, 2026
dsh plugin --profile web add dsh-verified-progress