DSH Plugins Marketplace

DSH Plugins

Plugins

/

dsh-local-perf

f

dsh-local-perf

Manifest valid

Durable DeepSeek Harness bundle: local-model performance tuning as a re-installable plugin layer (compaction, tool-result pruning, time context, cloud title routing, text-toolcall guard) — survives ds

hasBundlePatch

dsh-local-perf

Durable DeepSeek Harness plugin bundle carrying the local-model performance tuning process as a re-installable layer — so it survives dsh updates instead of living in hand-edited patch files that a version bump can wipe or whose rationale dies with the author.

Install once, re-apply forever:

# from the GitHub repository
git clone https://github.com/flowingboy/dsh-local-perf.git /tmp/dsh-local-perf
dsh plugin --profile web add file:/tmp/dsh-local-perf

# or from a checkout: dsh plugin --profile web add file:./dsh-local-perf

The bundle joins the profile's layer stack after dsh-base / dsh-web-app, re-applies all tuning rows on every boot, and vendors its own copy of the text-toolcall-guard plugin (self-contained, no dsh checkout required).


The complete performance process (why every knob exists)

Everything below was learned the hard way on an M5 Max running four local OpenAI-compatible servers (Ollama / MLX / Rapid-MLX / MLX-DSpark) against the DSH web GUI. The bundle encodes the conclusions; this README preserves the reasoning.

Incident log

DateSymptomRoot causeFix
2026-08-18"fan spin / no response"A 114,650-token prefill on Qwen3.5-122B — session context had grown unbounded because the web bundle disables auto-compactionRe-enable compaction-basic at thresholdRatio: 0.6
2026-08-19session-title starves the interactive stepThe title LLM request fires in the same second as a turn's first step; on the single-slot mlx-dspark server it queues behind the interactive generation and one of them starves past the idle budgetRoute session-title-llm to the cloud model (deepseek-official / deepseek-v4-flash)
2026-08-19mlxdspark timeouts with tools presentThe mlx-dspark server buffers the whole generation and emits no data events until it finishes; its 15s SSE keepalive comments are discarded by the OpenAI SDK parser and never reset the idle watchdogRaise timeoutMs / streamIdleTimeoutMs to 600000
2026-08-20tool calls appear as literal textLocal Qwen3.8-27B-8bit fell out of the structured tool_calls protocol under long tool-heavy steps and wrote <tool_call> prose the harness never executesShip the text-toolcall-guard plugin (vendored here)
recurringhallucinated "today"No clock context in the promptEnable time-context (Asia/Shanghai, 10 min refresh)

Layer 1 — model configs (settings.example.yaml → ~/.dsh/settings.yaml)

Machine-specific (paths, ports, model ids), so the bundle carries them as a template, not a runtime patch. Copy the llm-pi-ai section into ~/.dsh/settings.yaml on a fresh machine.

The recurring principles:

  • Timeout ≠ prefill tolerance. A slow local server needs timeoutMs + streamIdleTimeoutMs ≥ worst-case prefill + reasoning + decode. Gemma 4 31B prefills at ~180 tok/s (system prompt + tool schemas ≈ 13k tokens → ~70s), mlx-dspark buffers whole generations, so both budgets sit at 300–600s.
  • Retry only TRANSPORT. Connection-level failures happen before prefill and are cheap; a TIMEOUT must never re-prefill a long prompt.
  • Context window ≤ practical prefill budget. 262144 tokens at ~180 tok/s is minutes of prefill. Lower to 32K–64K; compaction at 0.6× keeps sessions safely under the server limit.
  • maxTokens ≤ decode budget. At ~27 tok/s decode, 16K output is ~10 min. Cap at 8192–16384 so one step's worst case fits the timeout budgets.
  • Reasoning effort default "off" (or the server's lowest level) for quick, low-latency local loops; the local model's thinking stream still renders as a DSH reasoning block when enabled.

Layer 2 — cordis rows (cordis.patch.yml)

RowWhatWhy
time-contextper-step clockkills hallucinated dates
compaction-basicauto-compact at 60%bounds prefill; the 08-18 incident fix
tool-result-prunerdrop stale tool resultskeeps them off later requests
command-compactmanual /compactescape hatch
session-title-llmtitle via cloud modelkeeps the local slot free for the interactive step

Layer 3 — text-toolcall guard (plugins/text-toolcall-guard)

Vendored from @deepseek-ai/dsh-text-toolcall-guard (built lib/ + src/). When a step closes with no native tool calls but the assistant text carries <tool_call> / <function=…> markers at line starts, the guard steers a corrective message so the model re-issues the call natively. Bounded to maxCorrections (2) per turn per agent; marker detection requires line-start placement so prose that merely quotes the format is not corrected. Peers are resolved at runtime from the dsh installation's profiles/node_modules fallback (the designed out-of-tree-plugin path), so no registry fetch needed.


Installation

# from this directory
dsh plugin --profile web add file:$(pwd)
# verify the layer joined the stack
dsh --profile web --dump-config | grep -A3 "dsh-local-perf"

The web profile's cordis.patch.yml should then only hold rows this bundle does NOT own (currently: none — everything moved into the bundle).

Updating the bundle

git pull                       # or edit locally
cd plugins/text-toolcall-guard # rebuild the vendored guard if its src changed
pnpm exec tsc -p tsconfig.json --outDir lib --declarationDir lib/types
# reinstall the layer
dsh plugin --profile web add file:$(pwd)

Publishing (GitHub / dsh-plugin ecosystem)

Published at https://github.com/flowingboy/dsh-local-perf (public, main, topics: dsh-plugin). To re-publish after local edits:

git add -A && git commit -m "dsh-local-perf: ..."
git push origin main
# topics (once)
gh repo edit dsh-local-perf --add-topic dsh-plugin

Layout

cordis.patch.yml                  the perf layer (all tuning rows)
settings.example.yaml             model-config template (copy to ~/.dsh/settings.yaml)
plugins/text-toolcall-guard/      vendored guard plugin (lib + src)
README.md                         this document — the preserved process

Comments

Loading…

Similar plugins

dsh-context-budget

by d3vmeh

DeepSeek Harness plugin: keep a local model's context at a size your GPU handles well (measured prefill speed, hard ceiling, early compaction)

Models & ProvidersTools & CapabilitiesManifest valid

★ 0

↓ 74/wk

MIT

JavaScript

Aug 30, 2026

dsh plugin --profile web add dsh-context-budget

by Apageoflove

Local-first experiment and evaluation workbench plugin for DeepSeek Harness (DSH).

Terminal & ClientsDevelopment & InfrastructureManifest valid

★ 3

↓ 5/wk

MIT

JavaScript

Sep 18, 2026

dsh plugin --profile web add dsh-arena

by ZK-Andy

Continual self-evolution plugin for DeepSeek Harness: versioned, auditable, rollback-safe harness state refined from session trajectories, with a benchmark-driven validation loop.

Tools & CapabilitiesDevelopment & InfrastructureWorkflow & AutomationSecurity & AuditManifest valid

★ 19

↓ 2.3k/wk

MIT

TypeScript

Sep 30, 2026

dsh plugin --profile agent add dsh-continual-evolve

by Missher12

Privacy-bounded self-improvement plugin for DeepSeek Harness

Workflow & AutomationTerminal & ClientsManifest valid

★ 0

MIT

TypeScript

Sep 8, 2026

dsh plugin --profile web add dsh-missher-evolution

by chuxindd

Persistent task state and context-aware compaction for DeepSeek Harness coding sessions.

Terminal & ClientsManifest valid

★ 0

MIT

TypeScript

Sep 20, 2026

dsh plugin --profile web add dsh-context-enhancement

by Electricitysheep

Per-round reasoning_effort optimizer for DeepSeek Harness (dsh): auto-downgrades tool-call reasoning for simple tool chains, lifting back for heavy work. Cuts thinking time between tool calls.

Tools & CapabilitiesManifest valid

★ 9

TypeScript

Aug 17, 2026

dsh plugin --profile web add dsh-tool-turbo