dsh-plugin-codemode
Manifest valid★ 1Pi-style Code Mode for DeepSeek Harness (DSH): Programmatic Tool Calling via QuickJS-WASM sandbox
dsh-plugin-codemode
English | 中文说明
Pi-style Code Mode (Programmatic Tool Calling) for DeepSeek Harness (DSH).
Turn verbose O(N) multi-turn ReAct loops into a single O(1) JavaScript orchestration script executed in an isolated memory sandbox.
⚡ The Benchmark: Hard Numbers
Tested against all 17 real markdown documents in the repository's docs/ folder (including a 131 KB operational ledger, totaling 215,011 raw characters). The task: scan all documents, filter by keywords (JEV, Polymarket, Architecture), extract titles & line metrics, and return a sorted Top-5 summary JSON.
| Metric | Traditional ReAct Loop | Code Mode (Sandbox Orchestration) | Measured Improvement |
|---|---|---|---|
| Data Loaded into Context | 215,011 chars | 676 chars | 99.7% Reduction (318:1 distillation) |
| Estimated Input Tokens | ~67,191 tokens | ~211 tokens | 99.7% Saved (67k tokens spared) |
| Agent / LLM Turns | 19 round trips | 1 turn | 94.7% Fewer Turns (18 turns eliminated) |
| Local I/O & Compute Time | 3.94 ms | 2.76 ms | Parity (sub-3ms execution) |
| End-to-End Elapsed Time | ~47.5 s (at 2.5s / LLM turn) | ~2.5 s | 19× Faster (94.7% latency drop) |
| Computation Accuracy | Prone to LLM counting hallucinations | Native JS Array.filter().sort() | 100% Deterministic & Accurate |
Reproduce locally:
node tests/benchmark-codemode-vs-react.mjs
💡 Why Code Mode?
Modern LLM agents suffer from tool explosion and context pollution:
- Context Bloat: Reading 20 files fills the transcript with 200k tokens of raw data. The model loses track, forgets system instructions, or triggers expensive context compaction.
- High Latency: Every single tool call requires a full round trip of network latency and model inference.
- Calculation Flaws: Asking LLMs to count lines, sort arrays, or calculate statistics over huge text chunks often yields subtle hallucinations.
The Code Mode Paradigm
"Put deterministic things in code, non-deterministic in LLM." — Hacker News Community Consensus on Code Mode
Inspired by Pi 1.0 (Earendil)'s Code Mode architecture and Cloudflare's production agent rewrite:
- Pi.dev 1.0 Official Codemode Docs: Established the Programmatic Tool Calling paradigm via memory-isolated sandboxing and deferred tool exposure.
- Cloudflare / CamelAI Case Study: Rewrote agents from heavy VM containers to Pi Code Mode in Durable Objects, reporting an order-of-magnitude reduction in latency and token costs.
- Hacker News Discussion (#49019301): Highlighted that orchestrating multi-tool workflows via local sandboxed code delivers up to a 99.2% cost reduction compared to traditional ReAct loops.
With Code Mode, the agent writes a concise JavaScript async function. The script executes inside an isolated sandbox, concurrently calls registered host and MCP tools, filters out unnecessary data in memory, and only returns the final distilled result to the conversation context.
🏗 Architecture
┌──────────────────────────────────────────────────────────┐
│ DeepSeek Harness (DSH) │
│ │
│ ┌──────────────────────┐ ┌──────────────────────┐ │
│ │ LLM Conversation │ │ ctx.tools / MCP │ │
│ │ (Main Context) │ │ (read, pwsh, mcp..) │ │
│ └──────────┬───────────┘ └──────────▲───────────┘ │
│ │ script │ │
│ ▼ │ invoke │
│ ┌─────────────────────────────────────────┴──────────┐ │
│ │ dsh-plugin-codemode (Cordis Extension) │ │
│ │ │ │
│ │ ┌──────────────────────────────────────────────┐ │ │
│ │ │ Dual-Engine Sandbox (Isolated Context) │ │ │
│ │ │ │ │ │
│ │ │ • V8 VM Sandbox (Default): native speed, │ │ │
│ │ │ zero memory caps, unconstrained concurrency│ │ │
│ │ │ • QuickJS-WASM: WebAssembly memory sandbox │ │ │
│ │ │ │ │ │
│ │ │ - Proxy bridge: tools.<name>(args) │ │ │
│ │ │ - Supports Promise.all concurrency │ │ │
│ │ │ - Hard timeout guard (default 60s) │ │ │
│ │ │ - Recursion & leak prevention │ │ │
│ │ └──────────────────────────────────────────────┘ │ │
│ └──────────────────────────┬─────────────────────────┘ │
│ │ distilled output only │
│ ▼ │
│ ┌────────────────────────────────────────────────────┐ │
│ │ Return to LLM: [Logs] + [Return Value] │ │
│ └────────────────────────────────────────────────────┘ │
└──────────────────────────────────────────────────────────┘
✨ Features
- 🚀 Dual-Engine Execution:
- V8 VM Engine (
engine: 'vm', Default): Aligns with DSH's official PTC workflow architecture. Zero overhead, handles multi-megabyte payloads, native microsecond performance. - QuickJS-WASM Engine (
engine: 'quickjs'): Strict WebAssembly memory isolation, zero Node/OS footprint.
- V8 VM Engine (
- 🔄 Universal Tool Proxy & Dynamic Introspection:
- Access all DSH built-in tools (
read,edit,pwsh,glob) and MCP tools (mcp__github__*,mcp__cua*,mcp__tavily__*) viatools.<tool_name>(args). - Zero-Token Tool Discovery: Introspect tools dynamically at runtime without polluting system prompts or wasting context tokens:
tools.list(): Returns the full array of registered tool names in memory.tools.help('namePattern'): Inspects parameters and JSON Schemas on-demand inside the sandbox, cutting system prompt bloat by over 90%.
- Access all DSH built-in tools (
- 🛡 Hard-Safety Protections:
- Configurable hard timeout (default 60,000 ms) kills infinite loops automatically.
- Recursion blocker prevents scripts from invoking
codemodeinsidecodemode. - Result truncator caps output at 50,000 characters to prevent accidental prompt floods.
- 📦 Zero-Patch, 100% Reversible: Pure dynamic Cordis extension. Modifies zero core DSH runtime files.
📖 Example Scripts
1. Parallel File Scan & Distillation
// Scan 10 files simultaneously and extract matching headers
const files = await tools.glob({ path: 'src', pattern: '**/*.ts' });
const results = await Promise.all(files.slice(0, 10).map(async (file) => {
const code = await tools.read({ file_path: file });
const exportedFns = code.match(/export (?:async )?function \w+/g) || [];
return { file, count: exportedFns.length, functions: exportedFns };
}));
// Sort by number of exports and return top 3
return results.sort((a, b) => b.count - a.count).slice(0, 3);
2. Zero-Token Dynamic Introspection & On-Demand MCP Execution
// Discover tools dynamically without polluting the prompt with 140+ tool schemas
const available = tools.list();
console.log(`Discovered ${available.length} active tools`);
// Inspect parameters on-demand only when needed
const issueHelp = tools.help('github__list_issues');
console.log(issueHelp);
// Fetch latest issues via GitHub MCP and summarize in memory
const issues = await tools.mcp__github__list_issues({ owner: 'deepseek-ai', repo: 'deepseek-harness' });
const topIssues = issues.slice(0, 3).map(i => ({
number: i.number,
title: i.title,
author: i.user?.login
}));
return topIssues;
🚀 Installation & Setup
Option A: dsh plugin add (Recommended)
dsh plugin add dsh-plugin-codemode
The package declares dsh.bundle.patch (cordis.patch.yml), so the host row is inserted
into your profile automatically — nothing to hand-edit. Restart DSH and the codemode
tool appears.
Option B: Install from npm manually
In your DSH profile directory (e.g. ~/.dsh/profiles/web):
pnpm add dsh-plugin-codemode
# or: npm install dsh-plugin-codemode
Then add the row yourself to ~/.dsh/profiles/web/cordis.patch.yml:
- insert:
- id: plugin-codemode
name: 'dsh-plugin-codemode'
config: # 全部可选;省略即用默认值
toolName: 'codemode'
engine: 'vm' # 'vm' (default) or 'quickjs'
maxResultChars: 50000
timeoutMs: 60000
injectGuidance: true
Option C: Install from Source (Development)
git clone https://github.com/Yum-wu/dsh-plugin-codemode.git
cd dsh-plugin-codemode
npm install
npm run build
Link into your DSH web profile (~/.dsh/profiles/web/package.json):
{
"dependencies": {
"dsh-plugin-codemode": "link:C:/Users/Yum/Desktop/dsh-plugin-codemode"
}
}
Then add the row from Option B.
Restart DSH
Restart the DSH service from your desktop management console. The agent will immediately receive the codemode tool declaration and execution guidance.
⚙️ Configuration Reference
| Option | Type | Default | Description |
|---|---|---|---|
toolName | string | 'codemode' | Name of the tool declared to the LLM. |
engine | 'vm' | 'quickjs' | 'vm' | Execution engine: high-throughput V8 VM or WASM sandbox. |
maxResultChars | number | 50000 | Maximum length of distilled output returned to context. |
timeoutMs | number | 60000 | Hard deadline per script before auto-termination. |
injectGuidance | boolean | true | Injects Code Mode orchestration tips into system prompt. |
autoReasoning | boolean | false | Enables Auto Reasoning Effort takeover. Must be turned on explicitly — see below. |
🧠 Auto Reasoning Effort
Set reasoningEffort to the sentinel auto in your model config and turn on
autoReasoning: true in this plugin's config; the plugin then projects a legal tier based on
task complexity onto the model's real effort ladder before the request goes out.
# cordis.patch.yml
- id: agent-default-model
name: "@deepseek-ai/dsh-agent-default-model"
config:
provider: opencodex
model: google-antigravity/gemini-3.8-flash
reasoningEffort: auto # <- sentinel
- id: plugin-codemode
name: 'dsh-plugin-codemode'
config:
autoReasoning: true # <- without this the plugin never intervenes
Both are required: auto alone is handed to the host and raises UNSUPPORTED_REASONING_EFFORT;
autoReasoning alone does nothing unless the incoming tier really is auto (a manually picked
tier, or one already stored in the session header, is always left untouched).
| Aspect | Detail |
|---|---|
| Hook | cordis agent/request waterfall, registered {global, prepend}; the return value goes straight into llm.prepareCall. |
Why not llm/stream | Its options is deepFreezed for agent-loop requests, and cordis next(x) ignores arguments — the tier cannot be changed there. |
| Ladder source | ctx.llm.resolveModelInfo(provider, model).reasoning.efforts — never hardcoded. |
| Scoring | Deterministic 1-10 keyword tiers (auto-reasoning.ts): concurrency/deadlock/risk = 9, derivation/refactor = 7, routine dev = 5, lookup = 2. |
| Fallback | No ladder or a failing resolveModelInfo -> omit reasoningEffort so the model default applies. auto is never handed back to the host. |
| Visibility | Read-only route GET /api/codemode.auto-effort?sessionId=<id> on the host's shared /api channel (same-origin auth applies); the composer-dock pill polls it every 3s. |
| Session isolation | Decisions are archived per sessionId (LRU, cap 50, AUTO_DECISION_CAP). agent/request is a global waterfall, so a single module-level slot made every conversation's pill show the same "most recent" record. The session id comes from payload.agent.id, falling back to ctx.get('agents').currentInitiator() — the real host dispatches agent/request with {turn, step, signal} only, so the initiator boundary is the path that resolves in practice (lastAgentSource reports which one fired). |
Showing Auto in the model picker
The picker's effort menu is built verbatim from the model's declared reasoningEfforts,
so auto has to be listed there for the entry to exist at all:
- id: google-antigravity/gemini-3.8-flash
reasoningEfforts:
auto: auto # <- adds the "auto" entry to the picker
low: low
medium: medium
high: high
Verified on dsh-0.2.0-rc.2 (2026-10-04): with auto: auto declared, the picker offers
Auto, selecting it sends reasoningEffort: 'auto' to agent/request, and the plugin
replaces it before the request leaves. The pill turns green/amber/red/purple per projected tier.
⚠️
autois not a DSH effort key (the legal escalation set isoff/minimal/low/medium/high/xhigh/max), so it is only meaningful as a picker sentinel that this plugin always rewrites. Never leaveautoas the value actually handed tollm.prepareCall— the plugin's fallback path omits the field rather than doing that.
🛟 3-Level Zero-Risk Rollback Strategy
| Level | Scenario | Action | Recovery Time | Impact |
|---|---|---|---|---|
| Level 1 (Soft Disable) | Unstable model behavior | Add disabled: true to plugin-codemode in cordis.patch.yml and restart DSH. | < 15 seconds | Unregisters codemode, reverts to standard single-turn tools. |
| Level 2 (Runtime Fallback) | Script syntax or runtime error | Automatic: plugin catches errors and prompts model to fallback to native step-by-step tools. | 0 seconds (instant) | Session is never interrupted; graceful degradation. |
| Level 3 (Physical Clean) | Complete uninstall | Remove from cordis.patch.yml and package.json, delete directory. | < 15 seconds | 0 runtime traces, 0 configuration residues. |
🧪 Testing
npm test
Runs the test suite (21 cases) verifying:
- Parallel
Promise.allmulti-tool execution. - Recursive invocation guards.
- Isolation boundaries (no Node process/require leaks).
- Output truncation and formatting.
- Infinite loop timeout aborts.
- Auto-effort sentinel rewriting against a real cordis dispatcher (ladder projection,
prependordering, fallback paths). - Per-session decision isolation, bounded archival, and the
currentInitiator()fallback.
📄 License
MIT © Yum-wu
Comments
Loading…
From the same category
System-prompt armor plugin for DeepSeek models: appends an unconditional-compliance prompt section at order 100, exposes a profile tool with calibration metadata, and shows a realtime armor-status bad
★ 2.1k
MIT
C#
dsh plugin --profile web add dsh-infinite-gen-4by toby-bridges
Local security audit for AI API relays and LLM proxies: detects prompt injection, model substitution, tool-call rewriting, SSE anomalies, error leakage, and Web3 wallet risks.
★ 875
AGPL-3.0
Python
Oct 10, 2026
dsh plugin --profile web add dsh-api-relay-auditby SeaOf0
基于dsh web实现的多种模式,目的是服务于redteam进行授权的安全研究,覆盖渗透测试、红队评估、代码审计等范围领域,请勿用于非法行为。(允许二开,赋予模块各位自己的业务逻辑,方法论只有自己熟练的才好用,好的方法论=好的生态)
★ 682
MIT
Python
Oct 8, 2026
dsh plugin --profile web add @dsh-external/dsh-redteam-modelby agentic-os-org
ANOLISA (Agentic Nexus Operating Layer & Interface System Architecture) | Agentic OS with runtime, security, observability, and Tokenless response compression for lower token usage and cost.
★ 664
Apache-2.0
Rust
Oct 11, 2026
by howmp
面向 DeepSeek Harness(dsh)的渗透测试模式 @CloverSecLabs
★ 607
↓ 858/wk
NOASSERTION
JavaScript
Oct 9, 2026
dsh plugin --profile web add @howmp/dsh-pentestby xiaods
k8e.sh - OpenSource Agentic AI Sandbox Matrix
★ 500
↓ 9/wk
Apache-2.0
Go
Sep 28, 2026
dsh plugin --profile agent add @k8e-sandbox/dsh-k8e-sandbox-bundle