DSH Plugins Marketplace

DSH Plugins

Plugins

/

dsh-local-model-supervisor

h

dsh-local-model-supervisor

Discovered

AI-written. Lazy-start a local llama.cpp model on demand and release RAM/VRAM when idle, with a DSH plugin. Written by AI: Local LLM "start only when needed, release when not in use" — on-demand cold start + automatic release when idle + fake context upper bound to guarantee compression

Machine translated

dsh-local-model-supervisor

中文 | English

Lazy-start a local llama.cpp model when DSH needs it, release RAM/VRAM when it doesn't.

本地大模型「用到才起、不用就放」:按需冷启动 + 空闲自动释放 + 一个假的上下文上限来兜住压缩。

[!IMPORTANT] AI-written code. Every file in this repository was produced by an AI coding agent, not by hand — see AI-DISCLOSURE.md before you run it. It was verified on a single Windows machine, with no security review and no test suite.

本项目由 AI 编写。 全部代码与文档均由 AI 编码助手产出,不是手写的 —— 运行前请先读 AI-DISCLOSURE.md。仅在作者的一台 Windows 机器上验证过,无安全审计、无测试套件。

A local-model supervisor for DeepSeek Harness. It occupies the ports your DSH local routes already point at and runs llama-server on private internal ports: the model is loaded on the first request and all RAM/VRAM is given back once it goes idle. Ships with a DSH plugin so the whole thing starts and stops with DSH.

DSH ──> 127.0.0.1:8080 (supervisor, always resident: ~30 MB RAM, 0 VRAM)
              │
              ├─ first inference request ──> 127.0.0.1:9080 (llama-server, started on demand)
              └─ 500s with no requests   ──> llama-server killed, VRAM released

1. Why this exists

DSH's local routes go through @deepseek-ai/dsh-llm-pi-ai, whose config is just api: openai-completions + baseURL — it is a plain HTTP client. It has no hook that could execute llama-server.exe. So in practice: the model is statically configured and always visible in the model picker, but selecting it while nothing is running just fails to connect, and nothing will ever start it for you.

The decision to lazy-start has to happen before the request reaches llama-server, so it can only be intercepted at the HTTP layer. This project makes the supervisor occupy the ports DSH is already talking to and moves the real llama-server to internal ports — so no DSH configuration has to change at all, and curl / Python / any other client gets the same behaviour for free.

2. Features

FeatureDescription
One unified model listDeclare any number of local models in a single list. Give a path and the supervisor manages the process; give a url and it just forwards to a server you already run (Ollama, LM Studio, vLLM, …). Two fields per entry is enough.
Ports allocated automaticallyThe i-th model takes publicBase + i; managed models take the next free port at or above internalBase, skipping anything claimed. Collisions are rejected at startup, never at spawn time.
Routes generated for youscripts/export-dsh-routes.ps1 turns the list into a ready-to-paste DSH providers: block, so importing a model does not mean hand-writing YAML.
Cold start on demandThe first inference request spawns llama-server, waits for health, then proxies. Clients never need to retry.
Idle releaseAfter idleTimeoutSec (default 500s) with no requests the process is killed and both RAM and VRAM are returned. The countdown starts when the last request ends, so long generations are never cut off.
Metadata never wakes the GPU/health and /v1/models are answered synthetically; /props and /metrics return 503 while unloaded. Otherwise a client polling /v1/models would pin the VRAM forever.
Single GPU slotRequesting a different model stops the current one first (one card cannot hold two).
Soft context ceiling ("fake overflow")A fake ceiling below the real window forces DSH to compact while redundancy still exists — see section 5.
Compaction grace windowAfter a rejection, requests are let through for a while, so the summarization replay cannot be blocked by the same ceiling.
Cold-start protectionA 1-token warm-up plus a settle delay, and one automatic restart-and-replay if the server still dies on its first large request.
Streaming preservedByte-for-byte pipe forwarding; SSE works untouched.
DSH pluginStarts/stops with DSH, plus /localmodel status | unload | start from any chat.

3. Quick start

3.1 Requirements

  • Node.js ≥ 18 (standard library only — no third-party dependencies)
  • llama.cpp's llama-server (CUDA / Vulkan / CPU all work; developed against a Windows CUDA build)
  • PowerShell 7+ (optional, only for the .ps1 helper scripts; the supervisor itself is pure Node)

3.2 Configure

cp supervisor/models.example.json supervisor/models.json

models is one unified list that accepts both kinds of local model source. All an entry needs is a name plus either path or url:

{
  "idleTimeoutSec": 500,                       // how long before a model is released
  "ports": { "publicBase": 8080, "internalBase": 9080 },

  "defaults": {                                // applied to every managed model
    "ctxSize": 32768,
    "cacheTypeK": "q8_0",
    "cacheTypeV": "q4_0",
    "softContextRatio": 0.75                   // softContextLimit = ctxSize * this
  },

  "llamaServer": {
    "bin": "C:\\path\\to\\llama-server.exe",
    "args": [ "--host", "127.0.0.1", "--threads", "8", "--flash-attn", "on", "--jinja", ... ]
  },

  "models": [
    { "name": "qwen-2bit", "path": "D:\\models\\Qwen-2bit.gguf" },
    { "name": "qwen-4bit", "path": "D:\\models\\Qwen-4bit.gguf", "ctxSize": 16384, "mtp": true },
    { "name": "ollama",    "url":  "http://127.0.0.1:11434/v1" },
    { "name": "lmstudio",  "url":  "http://127.0.0.1:1234/v1" }
  ]
}
EntryBehaviour
path — managedThe supervisor spawns llama-server for that file, cold-starts it on the first request and releases RAM/VRAM when idle.
url — externalAn already-running OpenAI-compatible server (Ollama, LM Studio, vLLM, another llama.cpp). The supervisor only forwards to it on its own public port; it is never started or stopped, so external entries get no idle release.

Ports are handed out for you. The i-th entry takes publicBase + i, and each managed entry takes the next free port at or above internalBase — skipping anything already claimed, so even a deliberately overlapping base pair (8080 / 8081) can never hand the same port to two things. Set publicPort / internalPort on an entry to override; genuine duplicates are rejected at startup with a clear error instead of failing later at spawn time.

Backwards compatible. common is still accepted as an alias for llamaServer, and the legacy explicit schema (publicPort / internalPort / model / args per entry) keeps working unchanged.

3.3 Generate the DSH routes

Declaring a model is only half the job — DSH also needs a route pointing at it, and hand-writing YAML for every model does not scale. Generate it from the same config instead:

pwsh -File scripts/export-dsh-routes.ps1                          # print to stdout
pwsh -File scripts/export-dsh-routes.ps1 -OutFile ..\routes.yml   # or write it to a file

It emits a ready-to-paste providers: block — one route per model, pointing at that model's supervisor public port, with contextWindow filled from ctxSize. Paste it under - id: llm-pi-ai → config: → providers: in your DSH profile patch, then restart DSH.

The generator reads the config through supervisor.js --dump-models, so the allocation and defaulting rules live in exactly one place and cannot drift from what the supervisor actually does.

3.4 Run it

# Foreground (Ctrl+C to stop)
node supervisor/supervisor.js supervisor/models.json

# Or in the background, via the PowerShell helpers
pwsh -File scripts/start-supervisor.ps1
pwsh -File scripts/supervisor-status.ps1     # what is loaded, and when it will be released
pwsh -File scripts/unload-local-models.ps1   # release everything now (supervisor stays up)
pwsh -File scripts/stop-supervisor.ps1       # stop the supervisor itself

Then point your DSH local route's baseURL at http://127.0.0.1:8080/v1.

3.5 Install the DSH plugin (optional, recommended)

The plugin makes the supervisor start and stop with DSH and adds a /localmodel command:

dsh plugin --profile desktop add link:<this repo>/plugin

Then set supervisorScript to the absolute path of your supervisor/supervisor.js in the plugin config (plugin/cordis.patch.yml, config: section). It is required — the plugin cannot guess where you cloned the repo.

A DSH restart is required for the plugin to load. After the restart you will see an apply() line in the plugin log.

4. Configuration reference

Supervisor (supervisor/models.json)

FieldDefaultDescription
idleTimeoutSec500Idle time before the model process is killed
checkIntervalSec5How often the reaper checks
startTimeoutSec300Cold-start wait limit
stopGraceSec20Wait limit when stopping a process
softGraceSec180Compaction grace window after a soft-ceiling rejection (section 5)
startSettleSec1.5Settle delay after the warm-up request
models[].name / key—Identifier for the entry (also the default alias)
models[].path—Managed: the .gguf file to serve
models[].url—External: an already-running OpenAI-compatible base URL (e.g. http://127.0.0.1:11434/v1)
models[].publicPortpublicBase + iPort clients connect to (owned by the supervisor)
models[].internalPortnext free ≥ internalBasePort llama-server listens on (managed entries only)
models[].aliaskey (or --alias)Model id reported by /v1/models
models[].ctxSizedefaults.ctxSizeContext window for managed entries
models[].cacheTypeK / cacheTypeVdefaults.*KV cache quantisation
models[].mtpoffAdd --spec-type draft-mtp (only pays off if the model has a NextN head)
models[].softContextLimitctxSize × softContextRatioFake context ceiling; 0 disables
models[].argsauto-builtAdvanced/legacy escape hatch: pass these to llama-server verbatim instead
ports.publicBase / ports.internalBase8080 / 9080Automatic port allocation bases
defaults.*see exampleField defaults for managed entries
llamaServer.bin / .args—How to launch llama-server (common is a legacy alias)

Control endpoints (on every public port)

EndpointPurpose
GET /__supervisor/statusJSON: per-model loaded state, idle seconds, time to release, context limits
POST /__supervisor/unloadStop every loaded model immediately

Plugin config

FieldDefaultDescription
supervisorScriptrequiredAbsolute path to supervisor/supervisor.js
supervisorConfigsibling models.jsonSupervisor config path
logFile~/.dsh/local-model-supervisor.logPlugin log
publicPorts[8080]Ports probed to adopt an existing supervisor
autoStarttrueStart the supervisor with DSH
unloadOnExittrueRelease loaded models when DSH exits
stopOnExitfalseAlso stop the supervisor process on exit
startTimeoutMs60000How long to wait for the supervisor to become ready

Chat commands

CommandPurpose
/localmodel / /localmodel statusShow supervisor and per-model status
/localmodel unloadRelease RAM/VRAM immediately
/localmodel startMake sure the supervisor is running

5. How the soft context ceiling ("fake overflow") works

The problem. Your local route declares a 32k window, while dsh-compaction-basic's default headroom is headroomTokens = 65536 — larger than the whole window, so W−O−B goes negative. That route therefore cannot resolve a trigger threshold and proactive compaction never runs; the session just grows to the hard limit. Worse, when it gets there, compaction has to replay "system prompt + old history" to write the summary, and that replay request may itself exceed the window — so compaction fails.

The fix. The supervisor enforces a fake ceiling below the real window:

real window  32768  ────────────────────────────────┐
fake ceiling 24576  ──────────────┐                 │  ← 8192 tokens of redundancy for compaction
                                  ↓
                    reply with llama.cpp's own overflow error
                                  ↓
         DSH classifies it as CONTEXT_WINDOW_EXCEEDED -> overflow recovery
              (that path bypasses the threshold policy, reduces the oldest
               balanced range, and retries)

Three things make this work:

  1. Exact counting. Before forwarding, the supervisor calls llama.cpp's POST /v1/chat/completions/input_tokens (~35 ms, and it still reports the true token count even when the request already exceeds n_ctx). If counting fails the request is forwarded — a safety net must never block legitimate traffic on its own.
  2. Byte-identical error. pi-ai's OVERFLOW_PATTERNS has a dedicated entry /exceeds the available context size/i (commented llama.cpp server), so reproducing the native payload verbatim is guaranteed to be recognised.
  3. Compaction grace window. The summarization request replays "system prompt + masked range", which — when a single turn added a lot of history — can still exceed the fake ceiling. So after a rejection the ceiling is relaxed for softGraceSec, letting that replay through. Without it the summary request would be blocked by our own ceiling and the session would deadlock.

The fake ceiling must stay above DSH's own proactive-compaction trigger, or it would also block the summarization replay.

6. Measured results

Test machine: RTX 4070 Ti 12 GB · Ryzen 5 9600X · 31 GB RAM · Windows · llama.cpp CUDA build.

ScenarioResult
Supervisor idle~30 MB RAM, 0 VRAM
Cold start, 9.15 GB 2-bit model5.7 – 8.8 s including load, correct answer
Cold start, 13.26 GB 4-bit model13.6 s
Idle release (4-bit loaded)9,686 MiB of VRAM returned
Request after a full release7.6 s to start again
/health, /v1/models, /propsmodel not woken, VRAM unchanged
26,012-token request (fake ceiling 24,576 / real 32,768)HTTP 400, byte-identical to the native overflow
Same request straight to the internal port (bypassing the supervisor)HTTP 200 — proving the 8k redundancy is real
Same request immediately after a rejection (simulating the compaction replay)HTTP 200 (grace window)
A small request afterwardsgrace cleared; the next oversized request triggers again
Cold start + a 26k requestllama-server crashed before the warm-up fix; stable 200 after

7. Notes and limitations

  • Supervisor mode is exclusive with running llama-server yourself (both want the same public port). scripts/start-supervisor.ps1 stops an existing llama-server first.
  • The fake ceiling depends on DSH's error classification, which lives in pi-ai (see section 5), not in this project. If upstream ever changes the wording, the fake ceiling silently degrades to "as if it were not there". The native compaction-basic + modelPolicies approach remains the correct fix then — and the two can be used together without conflict.
  • Single GPU slot: only one model runs at a time; asking for another stops the current one (and pays its cold start).
  • The supervisor takes no VRAM, but it is necessarily excluded from the idle reaper — it has to stay resident to be able to start anything on demand.
  • It listens on 127.0.0.1 only and has no authentication. Put a reverse proxy and auth in front of it if you expose it to a network.

Many models

Nothing is hard-coded to two models: the list takes any number, ports are allocated from publicBase / internalBase, the single-GPU-slot logic loops over every entry, and both the plugin and the helper scripts read the resulting port list out of the same config (via supervisor.js --dump-models) instead of keeping their own hardcoded list. Verified with four models (two managed + two external) on deliberately adjacent port bases.

Two things to plan for as the list grows:

  • Every switch between models costs a cold start. Only one model can be resident on a card that cannot hold two, so A → B stops A and loads B (5–15 s depending on size). A client that hops between models will pay that repeatedly — pin it to one model if you can.
  • Keep the port ranges clear of your other services. N models occupy publicBase … publicBase + N − 1, and the managed ones take the next free ports from internalBase. Make sure that does not collide with Ollama's 11434, LM Studio's 1234, DSH's own web port, and so on. Duplicates inside the config are rejected at startup with a clear error rather than failing later at spawn time.

If you genuinely want two models resident simultaneously, run two supervisor instances with separate configs and non-overlapping port ranges — each manages its own GPU slot. They still share one card, so this only helps when the models really do fit together.

8. Directory structure

.
├─ AI-DISCLOSURE.md             AI authorship notice — please read this first
├─ supervisor/
│   ├─ supervisor.js            the supervisor itself (Node, zero dependencies)
│   └─ models.example.json      config template -> copy to models.json
├─ plugin/
│   ├─ package.json             DSH plugin package (dsh.bundle.patch -> the yml below)
│   ├─ cordis.patch.yml         mount declaration + config (supervisorScript required)
│   └─ lib/index.js             plugin: lifecycle + /localmodel command
├─ scripts/                     optional PowerShell helpers
│   ├─ start-supervisor.ps1     start the supervisor
│   ├─ stop-supervisor.ps1      stop the supervisor
│   ├─ supervisor-status.ps1    status
│   ├─ unload-local-models.ps1  release everything now
│   ├─ export-dsh-routes.ps1   generate the DSH `providers:` block from models.json
│   ├─ supervisor-ports.ps1     shared helper: read the port list out of the config
│   ├─ install-autostart.ps1    start at logon (scheduled task; unnecessary once the plugin is used)
│   └─ uninstall-autostart.ps1  remove it
└─ LICENSE

9. DSH plugin implementation notes

It runs DSH's own Electron binary as Node. DSH runs inside Electron, so process.execPath is the application binary, not node. The plugin starts the supervisor with ELECTRON_RUN_AS_NODE=1, which runs that binary as a plain Node process — so no separate Node installation is required, and no external path is assumed:

spawn(process.execPath, [supervisorScript, config], {
  env: { ...process.env, ELECTRON_RUN_AS_NODE: '1' },
  windowsHide: true, stdio: ['ignore', logFd, logFd],
})

Probe before spawning, and never drag DSH down. On startup it asks /__supervisor/status first and adopts an existing supervisor instead of spawning a second one; a supervisor that fails to start, a command that fails to register, or a log file that cannot be written all just produce a log line — they never fail DSH startup.

License

MIT

Comments

Loading…

Similar plugins

dsh-plugin-local-model

by DZQJOKER

Local model plugin for DeepSeek Harness: manage GGUF models in settings, auto-launch llama.cpp on first message, and auto-unload after 5 minutes of idle time to free VRAM.

Models & ProvidersTerminal & ClientsManifest valid

★ 3

MIT

JavaScript

Sep 29, 2026

dsh plugin --profile web add dsh-plugin-local-model

by Lbunc

为DSH接入本地大模型能力:在「设置→插件」页一键启停本地 llama.cpp 大模型(双槽x双模态x双预设),卡片内配置、一条命令安装、自动注册,装完即用| Enable local large model capabilities for DSH: One-click start/stop for local llama.cpp models (dual‑slot × dual‑modal ×

Models & ProvidersManifest valid

★ 9

↓ 161/wk

MIT

JavaScript

Oct 1, 2026

dsh plugin --profile web add dsh-local-llm-controller

by DoctorxPriestess

A DSH plugin that manages local llama.cpp GGUF model lifecycles, automatically loading and unloading models on demand through an OpenAI-compatible gateway.

Terminal & ClientsManifest valid

★ 0

MIT

JavaScript

Sep 20, 2026

dsh plugin --profile web add dsh-llama-model-manager

by IYIcode

自动探测本地模型上下文 dsh-ctx-probe — DSH plugin: auto-sync llm-pi-ai contextWindow with the real runtime window of local llama.cpp/Ollama servers (probe n_ctx per request; tighten before overflow, widen when t

Manifest valid

★ 0

MIT

JavaScript

Sep 8, 2026

dsh plugin --profile web add dsh-ctx-probe

by PerryLink

Local-model (Ollama) integration for DeepSeek Harness: discover, pull, remove, and inspect local models, route requests to them by task type or keyword with automatic fallback to the cloud, and get a

Models & ProvidersManifest valid

★ 18

↓ 1.4k/wk

Apache-2.0

TypeScript

Oct 10, 2026

dsh plugin --profile agent add dsh-local-ai

by starefinger

LLM adapter plugin for locally deployed Qwen models behind a vLLM OpenAI-compatible endpoint, with per-model multimodal switch, fully configurable reasoning efforts, and a web settings page for editin

Models & ProvidersManifest valid

★ 4

↓ 488/wk

MIT

TypeScript

Sep 24, 2026

dsh plugin --profile web add dsh-llm-qwen-local