DSH Plugins Marketplace

DSH Plugins

Plugins

/

Security & Audit

/

dsh-plugin-codemode

Y

dsh-plugin-codemode

Manifest valid★ 1

Pi-style Code Mode for DeepSeek Harness (DSH): Programmatic Tool Calling via QuickJS-WASM sandbox

UI (client)hasBundlePatch

dsh-plugin-codemode

npm version License: MIT Platform: DeepSeek Harness Tests: Passing Engine: Dual (V8 VM + QuickJS)

English | 中文说明

Pi-style Code Mode (Programmatic Tool Calling) for DeepSeek Harness (DSH).
Turn verbose O(N) multi-turn ReAct loops into a single O(1) JavaScript orchestration script executed in an isolated memory sandbox.


⚡ The Benchmark: Hard Numbers

Tested against all 17 real markdown documents in the repository's docs/ folder (including a 131 KB operational ledger, totaling 215,011 raw characters). The task: scan all documents, filter by keywords (JEV, Polymarket, Architecture), extract titles & line metrics, and return a sorted Top-5 summary JSON.

MetricTraditional ReAct LoopCode Mode (Sandbox Orchestration)Measured Improvement
Data Loaded into Context215,011 chars676 chars99.7% Reduction (318:1 distillation)
Estimated Input Tokens~67,191 tokens~211 tokens99.7% Saved (67k tokens spared)
Agent / LLM Turns19 round trips1 turn94.7% Fewer Turns (18 turns eliminated)
Local I/O & Compute Time3.94 ms2.76 msParity (sub-3ms execution)
End-to-End Elapsed Time~47.5 s (at 2.5s / LLM turn)~2.5 s19× Faster (94.7% latency drop)
Computation AccuracyProne to LLM counting hallucinationsNative JS Array.filter().sort()100% Deterministic & Accurate

Reproduce locally: node tests/benchmark-codemode-vs-react.mjs


💡 Why Code Mode?

Modern LLM agents suffer from tool explosion and context pollution:

  1. Context Bloat: Reading 20 files fills the transcript with 200k tokens of raw data. The model loses track, forgets system instructions, or triggers expensive context compaction.
  2. High Latency: Every single tool call requires a full round trip of network latency and model inference.
  3. Calculation Flaws: Asking LLMs to count lines, sort arrays, or calculate statistics over huge text chunks often yields subtle hallucinations.

The Code Mode Paradigm

"Put deterministic things in code, non-deterministic in LLM." — Hacker News Community Consensus on Code Mode

Inspired by Pi 1.0 (Earendil)'s Code Mode architecture and Cloudflare's production agent rewrite:

  • Pi.dev 1.0 Official Codemode Docs: Established the Programmatic Tool Calling paradigm via memory-isolated sandboxing and deferred tool exposure.
  • Cloudflare / CamelAI Case Study: Rewrote agents from heavy VM containers to Pi Code Mode in Durable Objects, reporting an order-of-magnitude reduction in latency and token costs.
  • Hacker News Discussion (#49019301): Highlighted that orchestrating multi-tool workflows via local sandboxed code delivers up to a 99.2% cost reduction compared to traditional ReAct loops.

With Code Mode, the agent writes a concise JavaScript async function. The script executes inside an isolated sandbox, concurrently calls registered host and MCP tools, filters out unnecessary data in memory, and only returns the final distilled result to the conversation context.


🏗 Architecture

┌──────────────────────────────────────────────────────────┐
│                   DeepSeek Harness (DSH)                 │
│                                                          │
│  ┌──────────────────────┐      ┌──────────────────────┐  │
│  │ LLM Conversation     │      │   ctx.tools / MCP    │  │
│  │ (Main Context)       │      │  (read, pwsh, mcp..) │  │
│  └──────────┬───────────┘      └──────────▲───────────┘  │
│             │ script                       │             │
│             ▼                              │ invoke      │
│  ┌─────────────────────────────────────────┴──────────┐  │
│  │ dsh-plugin-codemode (Cordis Extension)             │  │
│  │                                                    │  │
│  │  ┌──────────────────────────────────────────────┐  │  │
│  │  │ Dual-Engine Sandbox (Isolated Context)       │  │  │
│  │  │                                              │  │  │
│  │  │  • V8 VM Sandbox (Default): native speed,    │  │  │
│  │  │    zero memory caps, unconstrained concurrency│  │  │
│  │  │  • QuickJS-WASM: WebAssembly memory sandbox   │  │  │
│  │  │                                              │  │  │
│  │  │  - Proxy bridge: tools.<name>(args)          │  │  │
│  │  │  - Supports Promise.all concurrency          │  │  │
│  │  │  - Hard timeout guard (default 60s)          │  │  │
│  │  │  - Recursion & leak prevention               │  │  │
│  │  └──────────────────────────────────────────────┘  │  │
│  └──────────────────────────┬─────────────────────────┘  │
│                             │ distilled output only      │
│                             ▼                            │
│  ┌────────────────────────────────────────────────────┐  │
│  │ Return to LLM: [Logs] + [Return Value]             │  │
│  └────────────────────────────────────────────────────┘  │
└──────────────────────────────────────────────────────────┘

✨ Features

  • 🚀 Dual-Engine Execution:
    • V8 VM Engine (engine: 'vm', Default): Aligns with DSH's official PTC workflow architecture. Zero overhead, handles multi-megabyte payloads, native microsecond performance.
    • QuickJS-WASM Engine (engine: 'quickjs'): Strict WebAssembly memory isolation, zero Node/OS footprint.
  • 🔄 Universal Tool Proxy & Dynamic Introspection:
    • Access all DSH built-in tools (read, edit, pwsh, glob) and MCP tools (mcp__github__*, mcp__cua*, mcp__tavily__*) via tools.<tool_name>(args).
    • Zero-Token Tool Discovery: Introspect tools dynamically at runtime without polluting system prompts or wasting context tokens:
      • tools.list(): Returns the full array of registered tool names in memory.
      • tools.help('namePattern'): Inspects parameters and JSON Schemas on-demand inside the sandbox, cutting system prompt bloat by over 90%.
  • 🛡 Hard-Safety Protections:
    • Configurable hard timeout (default 60,000 ms) kills infinite loops automatically.
    • Recursion blocker prevents scripts from invoking codemode inside codemode.
    • Result truncator caps output at 50,000 characters to prevent accidental prompt floods.
  • 📦 Zero-Patch, 100% Reversible: Pure dynamic Cordis extension. Modifies zero core DSH runtime files.

📖 Example Scripts

1. Parallel File Scan & Distillation

// Scan 10 files simultaneously and extract matching headers
const files = await tools.glob({ path: 'src', pattern: '**/*.ts' });
const results = await Promise.all(files.slice(0, 10).map(async (file) => {
  const code = await tools.read({ file_path: file });
  const exportedFns = code.match(/export (?:async )?function \w+/g) || [];
  return { file, count: exportedFns.length, functions: exportedFns };
}));

// Sort by number of exports and return top 3
return results.sort((a, b) => b.count - a.count).slice(0, 3);

2. Zero-Token Dynamic Introspection & On-Demand MCP Execution

// Discover tools dynamically without polluting the prompt with 140+ tool schemas
const available = tools.list();
console.log(`Discovered ${available.length} active tools`);

// Inspect parameters on-demand only when needed
const issueHelp = tools.help('github__list_issues');
console.log(issueHelp);

// Fetch latest issues via GitHub MCP and summarize in memory
const issues = await tools.mcp__github__list_issues({ owner: 'deepseek-ai', repo: 'deepseek-harness' });
const topIssues = issues.slice(0, 3).map(i => ({
  number: i.number,
  title: i.title,
  author: i.user?.login
}));

return topIssues;

🚀 Installation & Setup

Option A: dsh plugin add (Recommended)

dsh plugin add dsh-plugin-codemode

The package declares dsh.bundle.patch (cordis.patch.yml), so the host row is inserted into your profile automatically — nothing to hand-edit. Restart DSH and the codemode tool appears.

Option B: Install from npm manually

In your DSH profile directory (e.g. ~/.dsh/profiles/web):

pnpm add dsh-plugin-codemode
# or: npm install dsh-plugin-codemode

Then add the row yourself to ~/.dsh/profiles/web/cordis.patch.yml:

- insert:
    - id: plugin-codemode
      name: 'dsh-plugin-codemode'
      config:                  # 全部可选;省略即用默认值
        toolName: 'codemode'
        engine: 'vm'           # 'vm' (default) or 'quickjs'
        maxResultChars: 50000
        timeoutMs: 60000
        injectGuidance: true

Option C: Install from Source (Development)

git clone https://github.com/Yum-wu/dsh-plugin-codemode.git
cd dsh-plugin-codemode
npm install
npm run build

Link into your DSH web profile (~/.dsh/profiles/web/package.json):

{
  "dependencies": {
    "dsh-plugin-codemode": "link:C:/Users/Yum/Desktop/dsh-plugin-codemode"
  }
}

Then add the row from Option B.

Restart DSH

Restart the DSH service from your desktop management console. The agent will immediately receive the codemode tool declaration and execution guidance.


⚙️ Configuration Reference

OptionTypeDefaultDescription
toolNamestring'codemode'Name of the tool declared to the LLM.
engine'vm' | 'quickjs''vm'Execution engine: high-throughput V8 VM or WASM sandbox.
maxResultCharsnumber50000Maximum length of distilled output returned to context.
timeoutMsnumber60000Hard deadline per script before auto-termination.
injectGuidancebooleantrueInjects Code Mode orchestration tips into system prompt.
autoReasoningbooleanfalseEnables Auto Reasoning Effort takeover. Must be turned on explicitly — see below.

🧠 Auto Reasoning Effort

Set reasoningEffort to the sentinel auto in your model config and turn on autoReasoning: true in this plugin's config; the plugin then projects a legal tier based on task complexity onto the model's real effort ladder before the request goes out.

# cordis.patch.yml
- id: agent-default-model
  name: "@deepseek-ai/dsh-agent-default-model"
  config:
    provider: opencodex
    model: google-antigravity/gemini-3.8-flash
    reasoningEffort: auto          # <- sentinel
- id: plugin-codemode
  name: 'dsh-plugin-codemode'
  config:
    autoReasoning: true            # <- without this the plugin never intervenes

Both are required: auto alone is handed to the host and raises UNSUPPORTED_REASONING_EFFORT; autoReasoning alone does nothing unless the incoming tier really is auto (a manually picked tier, or one already stored in the session header, is always left untouched).

AspectDetail
Hookcordis agent/request waterfall, registered {global, prepend}; the return value goes straight into llm.prepareCall.
Why not llm/streamIts options is deepFreezed for agent-loop requests, and cordis next(x) ignores arguments — the tier cannot be changed there.
Ladder sourcectx.llm.resolveModelInfo(provider, model).reasoning.efforts — never hardcoded.
ScoringDeterministic 1-10 keyword tiers (auto-reasoning.ts): concurrency/deadlock/risk = 9, derivation/refactor = 7, routine dev = 5, lookup = 2.
FallbackNo ladder or a failing resolveModelInfo -> omit reasoningEffort so the model default applies. auto is never handed back to the host.
VisibilityRead-only route GET /api/codemode.auto-effort?sessionId=<id> on the host's shared /api channel (same-origin auth applies); the composer-dock pill polls it every 3s.
Session isolationDecisions are archived per sessionId (LRU, cap 50, AUTO_DECISION_CAP). agent/request is a global waterfall, so a single module-level slot made every conversation's pill show the same "most recent" record. The session id comes from payload.agent.id, falling back to ctx.get('agents').currentInitiator() — the real host dispatches agent/request with {turn, step, signal} only, so the initiator boundary is the path that resolves in practice (lastAgentSource reports which one fired).

Showing Auto in the model picker

The picker's effort menu is built verbatim from the model's declared reasoningEfforts, so auto has to be listed there for the entry to exist at all:

- id: google-antigravity/gemini-3.8-flash
  reasoningEfforts:
    auto: auto            # <- adds the "auto" entry to the picker
    low: low
    medium: medium
    high: high

Verified on dsh-0.2.0-rc.2 (2026-10-04): with auto: auto declared, the picker offers Auto, selecting it sends reasoningEffort: 'auto' to agent/request, and the plugin replaces it before the request leaves. The pill turns green/amber/red/purple per projected tier.

⚠️ auto is not a DSH effort key (the legal escalation set is off/minimal/low/medium/high/xhigh/max), so it is only meaningful as a picker sentinel that this plugin always rewrites. Never leave auto as the value actually handed to llm.prepareCall — the plugin's fallback path omits the field rather than doing that.


🛟 3-Level Zero-Risk Rollback Strategy

LevelScenarioActionRecovery TimeImpact
Level 1 (Soft Disable)Unstable model behaviorAdd disabled: true to plugin-codemode in cordis.patch.yml and restart DSH.< 15 secondsUnregisters codemode, reverts to standard single-turn tools.
Level 2 (Runtime Fallback)Script syntax or runtime errorAutomatic: plugin catches errors and prompts model to fallback to native step-by-step tools.0 seconds (instant)Session is never interrupted; graceful degradation.
Level 3 (Physical Clean)Complete uninstallRemove from cordis.patch.yml and package.json, delete directory.< 15 seconds0 runtime traces, 0 configuration residues.

🧪 Testing

npm test

Runs the test suite (21 cases) verifying:

  • Parallel Promise.all multi-tool execution.
  • Recursive invocation guards.
  • Isolation boundaries (no Node process/require leaks).
  • Output truncation and formatting.
  • Infinite loop timeout aborts.
  • Auto-effort sentinel rewriting against a real cordis dispatcher (ladder projection, prepend ordering, fallback paths).
  • Per-session decision isolation, bounded archival, and the currentInitiator() fallback.

📄 License

MIT © Yum-wu

Comments

Loading…

From the same category

dsh-infinite-gen-3

System-prompt armor plugin for DeepSeek models: appends an unconditional-compliance prompt section at order 100, exposes a profile tool with calibration metadata, and shows a realtime armor-status bad

Security & AuditManifest valid

★ 2.1k

MIT

C#

dsh plugin --profile web add dsh-infinite-gen-4

by toby-bridges

Local security audit for AI API relays and LLM proxies: detects prompt injection, model substitution, tool-call rewriting, SSE anomalies, error leakage, and Web3 wallet risks.

Security & AuditManifest valid

★ 875

AGPL-3.0

Python

Oct 10, 2026

dsh plugin --profile web add dsh-api-relay-audit

by SeaOf0

基于dsh web实现的多种模式,目的是服务于redteam进行授权的安全研究,覆盖渗透测试、红队评估、代码审计等范围领域,请勿用于非法行为。(允许二开,赋予模块各位自己的业务逻辑,方法论只有自己熟练的才好用,好的方法论=好的生态)

Security & AuditManifest valid

★ 682

MIT

Python

Oct 8, 2026

dsh plugin --profile web add @dsh-external/dsh-redteam-model

by agentic-os-org

ANOLISA (Agentic Nexus Operating Layer & Interface System Architecture) | Agentic OS with runtime, security, observability, and Tokenless response compression for lower token usage and cost.

Security & Audit

★ 664

Apache-2.0

Rust

Oct 11, 2026

Index only — not installable

by howmp

面向 DeepSeek Harness(dsh)的渗透测试模式 @CloverSecLabs

Security & AuditManifest valid

★ 607

↓ 858/wk

NOASSERTION

JavaScript

Oct 9, 2026

dsh plugin --profile web add @howmp/dsh-pentest

by xiaods

k8e.sh - OpenSource Agentic AI Sandbox Matrix

Security & AuditManifest valid

★ 500

↓ 9/wk

Apache-2.0

Go

Sep 28, 2026

dsh plugin --profile agent add @k8e-sandbox/dsh-k8e-sandbox-bundle