dsh-llm-verifier-pro
Manifest validLLM-as-a-Verifier plugin for DeepSeek Harness — fine-grained reward tools (verify_compare / verify_select / verify_track) with Probabilistic Pivot Tournament, plus a Best-of-N conversation mode with a Web settings panel.
dsh-llm-verifier-pro
A quality gate for DeepSeek Harness: sample N candidate answers, score them with fine-grained logprob rewards, and replay only the best one — instead of handing the model's first draft straight to you.
- Best-of-N conversation mode — each text turn is sampled N ways, ranked by
expected score, and the winner is replayed with a muted
⚡ Best-of-Nfooter; - Three
verify_*tools —verify_compare/verify_select/verify_trackfor self-checking, best-of-N selection, and progress tracking; - Paper method — fine-grained reward as the expectation over the verifier's
top-20 logprob distribution at the
<score_A>position (arXiv:2607.05391), with Bradley–Terry + Probabilistic Pivot Tournament (O(N·k), not O(N²)); - Fails open, always — sampling overruns, scoring failures, endpoints without logprobs: everything degrades gracefully, never a dead turn.
Based on the LLM-as-a-Verifier paper (arXiv:2607.05391);
engineering core from dsh-llm-as-a-verifier (TaurenMountain), product layer
from @aispin/plugin-verifier (Aispin), both MIT.
Installation
Requires a DeepSeek Harness profile
(web or headless). Add the plugin to the target profile — the compiled
lib/ is tracked in this repo, so GitHub installs work out of the box:
# from GitHub (recommended — works immediately)
dsh plugin --profile web add github:jmche/dsh-llm-verifier-pro
# or from a local checkout / path
dsh plugin --profile web add /path/to/dsh-llm-verifier-pro
The plugin registers as the bundle row llm-verifier-pro. Then configure it in
the profile's patch layer (see Configuration below), restart
dsh web, and the plugin exposes:
- three
verify_*tools to every agent; - the optional Best-of-N conversation mode (off by default).
Full walkthrough: docs/USER-GUIDE.md.
Three faces
1. Tools (agent calls them on demand)
verify_compare— fine-grained rewards (R_A, R_B) ∈ [0, 1] for one directed pairwise comparison under your criteria.verify_select— Probabilistic Pivot Tournament best-of-N selection: O(N·k) verifier comparisons instead of O(N²), seeded and reproducible.verify_track— per-step progress curve (A = 0% … T = 100%) over your trajectory, decoded from the logprob expectation.
2. Service (ctx.verifier)
ctx.verifier.verify / compare / select / track for code consumers.
3. Mode (Best-of-N conversation mode)
When enabled, every assistant turn that produces a final text answer is
sampled N ways and only the winning response is replayed to you. Tool-call
turns are never sampled — when the model's turn is an action (reading a
file, running a command, calling a tool), the turn is replayed exactly as
produced: Best-of-N ranks text answers only, and sampling a working turn
would waste tokens and produce unusable candidates. The decision is
re-evaluated per turn and is deliberately all-or-nothing — the mode covers
every conversation, there is no per-session tier. It is one switch: boN: true in the plugin config turns the mode on for every
conversation at boNCandidates; anything else is off.
Behavior change (dsh 0.1.7): there is no settings layer above the plugin config any more, so the plugin config is the only place the switch lives — see Configuration. (The
bo-nsession-preset tier is a separate, earlier removal — this plugin's own 0.2.0.)
Model mix (candidate diversity). Candidate 0 always rides the
conversation's own model (the greedy anchor). Each later slot draws a
{ provider, model } entry from boNModelMix in order; slots beyond the list
fall back to anchor-model variants at the sampling temperature. Configure it in
the patch layer.
provider is a REAL dsh provider route (omni-chat, omni-message,
deepseek-official…); model is the FULL model id exactly as that provider
advertises it (possibly containing its own /, e.g. agnes/agnes-2.5-flash).
A string entry is split at its FIRST / — so a model id that itself contains
/ (e.g. ollama-local/qwen3.8:27b) MUST be written with its real provider
(omni-chat/ollama-local/qwen3.8:27b); only a model id WITHOUT / can ride
the conversation's provider as a bare string.
boNModelMix:
- provider: omni-chat
model: agnes/agnes-2.5-flash
- provider: omni-message
model: opencode-go/minimax-m3
- provider: omni-chat
model: ollama-local/qwen3.8:27b
Every failed path fails open: a sampling overrun degrades Bo5 → Bo-K → a normal answer, with a muted footer explaining what happened. Never a dead turn.
Faithfulness to the paper
The tools, the service and the Bo-N mode implement the LLM-as-a-Verifier
method (arXiv:2607.05391) exactly as shipped by the
official repository
(MIT): fine-grained reward = expectation over the verifier's top-20 logprob
distribution at the <score_A> / <score_B> positions of a 20-letter
(A–T) scale, normalized to [0, 1] (paper Eq. 3.1); Bradley–Terry preference
p = σ(R_A − R_B) (Eq. 3.2); Probabilistic Pivot Tournament with a random
Hamiltonian-cycle ring pass, top-k pivots by mean preference and pivot
rounds — N + k(N−k) + C(k,2) = O(Nk) comparisons, seeded and reproducible
(Algorithm 1); criteria decomposition (C) and repeated evaluations (K);
per-checkpoint progress tracking (A = 0% … T = 100%); and the vLLM/SGLang
score-tag prefill pass for logit-restricted backends (paper Appendix B.6).
The pairwise and progress prompts match the official templates.
Documented deviations vs. the official repo / paper — none changes the method:
- Bo-N uses a full round-robin for N ≤ 3 (
src/bon.ts): PPT only applies from N ≥ 4, where it is cheaper and lower-variance; exhaustive scoring of tiny pools is exact.verify_selectalways uses PPT. - Defaults are weaker than the paper's headline protocol (G=20, K=8,
three-criterion decomposition):
verify_compareK=1,verify_selectK=4 (the official repo's default), the Bo-N turn K=1 with a singlecorrectnesscriterion. All are configurable (nEvaluations,criteria, …). verify_trackis the offline one-call variant (same as the officialtrack()): one call scores every checkpoint and sees the whole trajectory; the strict per-prefix protocol (the officialProgressTracker) is not ported.- No multimodal (image/video) inputs — the official repo accepts
images; the TS backend does not. - No persistent JSON score cache —
selectkeeps an in-memory cache per run only. - Not bit-reproducible across implementations — the PRNG is mulberry32
(not Python's
random), so a seed reproduces a tournament within JS but not the same ring as Python; criterion-id slugging uses-instead of_.
Configuration
Zero-config default: with no explicit baseUrl / apiKey / model, the
verifier follows the session —
same provider route, endpoint and model as the conversation. Turning on
Best-of-N alone gives the paper's self-verification experience: candidates
are sampled as variants of the conversation's own model, and that same model
grades them. The endpoint for the session provider is read from that adapter's
own Loader entry config (llm-pi-ai style, providers.<name>.baseURL), through
the configEditor service. Only when the session provider is unknown does the
resolution fall back to:
plugin config → session provider endpoint → OPENAI_BASE_URL /
OPENAI_API_KEY / DEEPSEEK_API_KEY → api.deepseek.com.
The verifier must sit on an endpoint that returns token-level logprobs
(vLLM, SGLang, OpenAI, DeepSeek, and modern Ollama all do; a plain gateway
that strips logprobs will not). Non-DeepSeek servers get the optional
vLLM/SGLang prefill pass so score tags land exactly at the label position.
# ~/.dsh/profiles/<profile>/cordis.patch.yml
- id: llm-verifier-pro
config:
# Leave baseUrl/apiKey/model empty to follow the session model.
# Set them explicitly to use a dedicated scoring endpoint:
baseUrl: https://your-gateway/v1
apiKey: credential:YOUR_API_KEY_ENV
model: opencode-go/deepseek-v4-flash
boN: false # master switch — this line is the only place it lives
boNCandidates: 5
samplingMode: parallel # rollouts per turn: 'parallel' (default) fires N at once; 'serial' waits one-at-a-time (safer when several candidates share one slow local model)
showFooter: true
Jev as the selector (optional)
selector: jev hands every pairwise comparison — Best-of-N ranking,
verify_select, verify_compare — to a System One endpoint
(TypeSafe Jev) instead of the LLM verifier. Jev
does not generate text: each comparison is one request with the task and both
responses as state and one ordered Score question per slot; the reward is the
expectation over the Score levels, normalized to [0, 1] like the logprob
expectation, so the tournament, slot swaps and Bradley–Terry aggregation are
unchanged. A comparison typically returns in under a second. verify_track
always stays on the LLM verifier; an LLM verifier config that cannot resolve
(e.g. a missing credential: key) is logged and does not block Jev. The
default selector: llm changes nothing.
The endpoint is provider-agnostic: jevBaseUrl is a …/v1 base (the plugin
appends /systemone) or the full …/systemone URL.
- id: llm-verifier-pro
config:
selector: jev
# TypeSafe directly (the default when jevBaseUrl/jevModel are empty):
# jevBaseUrl: https://api.typesafe.ai/v1
# jevModel: jev-latest
# jevApiKey: env:TYPESAFE_API_KEY
# OpenCode Zen (jev-1.13-free is free for a limited time):
jevBaseUrl: https://opencode.ai/zen/v1
jevModel: jev-1.13-free
jevApiKey: env:OPENCODE_API_KEY # credential:<name> | env:VAR | plain; empty sends no key
Jev's context budget is 32k tokens for the state plus the longest question;
a comparison whose two responses exceed it returns HTTP 400
(max_tokens_exceeded) and the turn fails open like any other verifier error.
Jev 1.13 is weak at arithmetic, counting and dates
(jaggedness notes),
and the state is sent to the configured provider — check its data policy.
Configuring in the Web UI
Open Plugins → dsh-llm-verifier-pro in dsh web: the bundle page carries a
form for every parameter below except apiKey, jevApiKey, deepseek, prefill and
compare/select/track. Save writes the same cordis.patch.yml block
shown above, and the change applies to the next turn without a restart (the
form's fields are .volatile() Config fields, which the Loader commits into
the running plugin). The excluded fields are edited in the patch; changing one
there remounts the plugin.
Out-of-range values (a timeout below 1 ms, fewer than 2 candidates, a
temperature outside 0–2, a fractional count) are refused on save. The form is
bound to the entry id llm-verifier-pro; a row renamed in the patch, or a
second instance under another id, is configured in the patch only. Model mix
and criteria take one entry per line.
Parameters
Every parameter lives in the plugin config (the profile's cordis.patch.yml)
and overrides the built-in default. Since dsh 0.1.7 that is the only layer:
~/.dsh/settings.yaml no longer exists (dsh imported it into the profile patch
and renamed it settings.yaml.imported).
| Parameter | Default | Meaning |
|---|---|---|
boN | false | Best-of-N master switch. |
boNCandidates | 5 | Candidates sampled per text-answer turn. |
samplingTemperature | 0.7 | Diversity temperature for the sampled candidates. |
samplingMode | parallel | Rollout schedule: parallel (all at once) or serial (one at a time). |
boNModelMix | [] | Model mix for non-anchor candidates; empty = same-model (follow the session). |
timeoutMs | 300000 | Per-request verifier HTTP timeout in ms. |
timeoutMsBoN | 300000 | Wall-clock budget for the sampling phase. |
verifyTimeoutMsBoN | 300000 | Wall-clock budget for the ranking phase. |
showFooter | true | Append the muted ⚡ Best-of-N … footer under the winner. |
criteria | [] | Extra grading criteria appended to the comparison prompt. |
boNPivots | 2 | PPT pivot count k. |
boNSeed | 0 | Seed for the tournament ring pass. |
verifier | '' | Verifier as a provider/model route; empty = follow the session model. |
autoDegrade | true | Fall back to sampling scoring when the endpoint lacks logprobs. |
baseUrl / apiKey / model | '' | Explicit verifier endpoint three-part; empty = follow the session. |
maxConcurrency | 8 | Max in-flight verifier calls. |
selector | llm | Pairwise scorer: llm (the verifier above) or jev (a System One endpoint). |
jevBaseUrl | TypeSafe | System One …/v1 base or full …/systemone URL; empty = https://api.typesafe.ai/v1. |
jevModel | jev-latest | System One model id, e.g. jev-1.13-free on OpenCode Zen. |
jevApiKey | '' | credential:<name>, env:VAR or plain; empty sends no Authorization header. |
deepseek | auto | Force the DeepSeek call path. |
prefill | true | vLLM/SGLang score-tag prefill pass. |
compare / select / track | true | Register the three verify_* tools. |
maxConcurrency, deepseek, prefill and compare/select/track are
deployment knobs rather than things a user tunes per turn.
Development
npm install
npm run check # typecheck + tests (112 tests)
npm run build # tsc
License
MIT. Implementation ports from the two upstreams above, both MIT.
Comments
Loading…
From the same category
by awesome-dsh-plugin
A curated list of plugins for DeepSeek Harness (dsh) · DeepSeek Harness 插件精选列表
★ 18.3k
CC0-1.0
Python
Oct 10, 2026
by 0xsline
DeepSeek Harness (DSH) ecosystem: curated plugins, tools, and infrastructure from dsh-external/hub and the public dsh-plugin topic.
★ 1.2k
CC0-1.0
Python
Oct 10, 2026
by pax-beehive
Open-source CLI, schemas, resolver, and DSH agent tools for DSH Plugin Hub
★ 450
MIT
TypeScript
Oct 6, 2026
by xiajiajun516
DeepSeek Harness (DSH) backup & restore plugin — export, import, migrate and sync your complete DSH configuration, plugins, MCP servers, skills and workspace. One-click migration to another machine.
★ 176
MIT
TypeScript
Oct 8, 2026
dsh plugin --profile web add dsh-config-managerby yjh051108
推荐组件(非必须):DeepSeek Harness 运行时注入器;已随 dsh-routing-suite 单仓库化保留,本仓库继续维护/发布。
★ 164
TypeScript
Sep 18, 2026
dsh plugin --profile web add @dsh-external/dsh-super-injectorby jigjoy-ai
A CLI that turns a goal into a pull request - and a sandbox for testing concurrent AI coding agents on the Mozaik runtime.
★ 124
MIT
TypeScript
Oct 2, 2026