DSH Plugins Marketplace

DSH Plugins

Plugins

/

dsh-agent-arena

L

dsh-agent-arena

Manifest valid2

Isolated multi-model coding matches with deterministic verification, scoring, and reports

UI (client)hasBundlePatch

dsh-agent-arena

English | 中文

A DSH coding arena for comparing 2–4 configured models in isolated Git worktrees, validating their changes deterministically, reviewing every diff, and explicitly applying one winner.

Screenshot

Agent Arena results and diff review

Generated with GPT Image from the implemented Client layout and feature set; runtime appearance follows the active DSH theme and viewport.

Execution and persistence

  • Creates detached worktrees under the operating system temporary directory, outside the compared repository.
  • Restricts recursive cleanup to the exact two-level match/contestant path under that dedicated Arena directory, including resolved junction and symlink checks.
  • Starts each contestant with the public ctx.agents.create, SessionId, createUserMessage, and Agent.followup APIs.
  • Executes Git and validation argv through ctx.subprocess.spawn; a shell is never used and shell operators are rejected.
  • Persists match reports through storageDomain. An interrupted nonterminal match is marked failed after Host restart.
  • Polls the generated agentArena Remote namespace in Settings and supports Start, Cancel, diff review, and Apply winner without browser globals.

Scoring and application

Each validation has a positive weight. A contestant score is the percentage of total validation weight that exits successfully; score ties use contestant id as a stable deterministic tie-breaker. No LLM judge is used. The form starts without a generic validation because git status --porcelain=v1 normally exits successfully regardless of whether it prints changes; add the repository's own test, typecheck, or build command before starting.

A match requires clean git status --porcelain=v1 before worktrees are created. The winner worktree remains available until explicit application. Apply repeats the cleanliness check, verifies that HEAD still equals the recorded base revision, runs git apply --check, then git apply --index --whitespace=error. Arena never auto-applies, commits, pushes, or rewrites history.

Validation command syntax

The Settings field accepts conservative whitespace-separated argv such as corepack pnpm test or corepack pnpm typecheck. Quotes, shell variables, pipes, redirections, command separators, backticks, and newlines are rejected. Configure commands whose arguments do not require shell quoting.

Install

dsh plugin --profile web add github:LeemanCheung/dsh-agent-arena

Restart the existing DSH Web process and refresh its page. The selected profile must provide agents, subprocess, storageDomain, Typert Remotes, Settings, and at least two usable provider/model routes..

Model Experience

Every contestant receives the user-entered Arena objective as one user message in its own session and worktree. Token and KV-cache usage therefore scales with 2–4 independent contestant sessions. Validation, scoring, comparison, cancellation, persistence, and application do not call another model.

Known limitations

Validation argv uses a deliberately conservative tokenizer and cannot represent arguments containing whitespace. Cancellation depends on the selected provider honoring the supplied abort signal. The Host captures binary-capable git diff HEAD, so staged, unstaged, and intent-to-add files are reviewed and applied consistently. Patches over the 1 MB review/apply bound are rejected before mutation rather than truncated.

Development

From the repository root run corepack pnpm typecheck, corepack pnpm test, corepack pnpm build, and corepack pnpm pack:check.

Version 1.0.1 is marked compatible with the DSH 0.1.2-rc.1 web profile after Windows QA loaded its Host service, browser Client, generated Remote namespace, and Settings section in both the isolated QA instance on port 3081 and the existing local profile on port 3080. The form showed no misleading prefilled validation and kept Start disabled until required fields were supplied. Automated coverage contains 15 passing tests. QA did not start a real model match, so provider execution and cancellation remain outside this compatibility claim. CI is configured to rebuild the committed lib artifacts on both Windows and Linux and reject any tracked or untracked difference.

MIT. See LICENSE.

Comments

Loading…

Similar plugins

dsh-multi-model-provider

by AlexKaiqi

Model catalog, portraits, Agent selection, and multimodal runtimes for DeepSeek Harness

Terminal & ClientsSessions & MessagesManifest valid

2

MIT

TypeScript

Sep 9, 2026

dsh plugin --profile web add dsh-multi-model-provider

by alison-xx

Visual workflows and multi-model evaluation for DeepSeek Harness

Manifest valid

0

MIT

TypeScript

Aug 14, 2026

dsh plugin --profile web add deepseek-harness-flow

by toolclub

Persistent multi-model workflow teams for DeepSeek Harness — dynamic lead planning, bounded DAGs, per-agent model/tools, Run Center and Token insights.

Workflow & AutomationTools & CapabilitiesModels & ProvidersManifest valid

176

MIT

TypeScript

Sep 19, 2026

dsh plugin --profile web add dsh-agent-team-gui

by lkshjd

DeepSeek Harness multi-agent debate plugin: isolated research, cross-examination, judge convergence (background job + resume + parallel waves)

Workflow & AutomationManifest valid

0

TypeScript

Aug 17, 2026

dsh plugin --profile web add @sky_sun/dsh-debate

by wowyuarm

Give DeepSeek Harness a persistent agent team for long-running collaboration

Workflow & AutomationManifest valid

32

755/wk

MIT

TypeScript

Sep 20, 2026

dsh plugin --profile web add @wowyuarm/dsh-agent-team

by changer-changer

A blind, fair, local DSH Web arena: same task, isolated worktrees, shared verification, judge before reveal.

Manifest valid

0

MIT

TypeScript

Aug 19, 2026

dsh plugin --profile web add dsh-blind-arena