DSH Plugins Marketplace

DSH Plugins

Plugins

/

dsh-tier-router

z

dsh-tier-router

Manifest valid

Automatic tier-based model routing for DeepSeek Harness (dsh): a virtual `smart` model classifies every request by difficulty (hard / normal / easy) and by vision need, then delegates it to the models you already configured.

UI (client)hasBundlePatch

dsh-tier-router

npm license dsh-plugin

Tier-based automatic model routing for DeepSeek Harness.

It registers one virtual model — Tier Router (auto) — that you select as your session model. Every request is then classified by difficulty (hard / normal / easy) and by whether it contains images, and delegated to the models you already configured in Settings → Models.

No extra upstream, no extra API key: the routing targets are your existing models — a local Codex route, a llama.cpp endpoint, an official API, anything.

中文说明 →

Why

Model pickers make you choose once and live with it: pick the strong model and you pay for it on "thanks, continue"; pick the cheap local model and it fumbles the hard refactor. dsh-tier-router makes that choice per request instead:

  • greetings, translations and short explanations → your easy tier (e.g. a free local model)
  • ordinary coding and file work → your normal tier
  • architecture, debugging, concurrency, migrations → your hard tier
  • anything with an image → your vision tier, with the image turned into structured text first

Features

  • Three-tier difficulty routing — heuristic classifier (default: zero cost, zero latency, deterministic), an optional LLM classifier, or Jev (TypeSafe System One), which answers the tier as a typed three-option choice instead of generating text. All three are cached.
  • Vision sidecar — when a request carries images, a vision model converts them into structured evidence (summary / OCR / layout) that replaces the image block as text; the difficulty tiers then answer. A legacy route mode sends the whole turn to the vision tier instead.
  • Ladder fallback — requested tier → the nearest remaining tiers → your default model. easy falls to normal before hard, so a small local model that cannot answer never hands a one-line question to the most expensive route; normal and hard keep escalating first.
  • Fail-open — an error finish chunk is only emitted when every route failed; requests are never silently swallowed.
  • No forks, no patches — pure adapter-level routing (ctx.llm.registerAdapter + prepareCall); the Harness source is untouched.
  • Built-in settings card — pick each tier from a live model catalog (annotated with image support and available reasoning efforts) in Settings → Tier Router, and watch route counters and recent failures.
  • Two-level reasoning effort — the chat selector (Off / High) is the master switch; when it is on, each tier uses the effort you configured for it.
  • Recursion-safe — a tier that points at the router itself is skipped, so routing can never loop.

Install

dsh plugin --profile web add dsh-tier-router

Then restart dsh web (host plugins load at boot). The package declares dsh.bundle.patch, so dsh plugin add appends it to the profile's layer stack and its own cordis.patch.yml inserts the plugin row — do not add a second row by hand.

Installing straight from GitHub also works (this package is plain ESM with no build step):

dsh plugin --profile web add github:zhangzhangco/dsh-tier-router

Updating. dsh plugin add records a caret range, and for a 0.x version a caret pins the minor: ^0.1.0 can never resolve to 0.2.0. Ask for the version (or the tag) explicitly:

dsh plugin --profile web add dsh-tier-router@latest

Right after a release, a local install can still resolve the previous version from pnpm's cached registry index; pinning the exact version (dsh-tier-router@0.2.0) always re-fetches.

Quick start

  1. Open the Web GUI → Settings → Tier Router and give each of the four tiers a provider + model. The pickers list every model you have configured, with ✓ Vision and reasoning-effort metadata.
  2. Pick Tier Router (auto) as the session model in the chat model selector.
  3. Send messages as usual. Route counters and recent failures are on the same settings page.

The defaults in this repository point at the author's own local routes (codex-local, gpudev). Set the four tiers to your own models. Tiers left empty are skipped automatically and the request falls back — nothing breaks.

How do I see which model was chosen?

Open Settings → Tier Router. Below the counters, Recent routing decisions lists the last 20 requests, newest first, each showing the model that answered, the classified tier, the estimated request size, and why that route won:

21:47:12  codex-local/gpt-5.6-terra  ·  hard  ·  ~181k tok  ·  by heuristic  ·  new turn  ·  hard tier  ·  why: 3 hard signal(s)
21:46:40  gpudev/qwen3.8-27b-q5  ·  easy  ·  ~2k tok  ·  by heuristic  ·  same turn  ·  easy tier  ·  why: social signal(s) in short message
21:45:03  codex-local/gpt-5.5  ·  normal  ·  ~181k tok  ·  by llm-cache  ·  same turn  ·  normal tier  ·  skipped (window too small) gpudev/qwen3.8-27b-q5 (131072 < est 181000, easy tier (fallback))

Three fields exist specifically to make a surprising tier explainable:

  • by <classifier> — which classifier produced the level: heuristic; llm, llm-cache, llm-timeout, llm-error, llm-unavailable; or jev, jev-cache, jev-low-confidence, jev-timeout, jev-auth, jev-rate-limit, jev-error, jev-unavailable. A tier that came from the heuristic because a semantic classifier timed out, lost its API key or abstained says so instead of being indistinguishable from a real classification.
  • why: … — the classifier's own one-sentence reason, or the heuristic's scoring reasons. The route reason (normal tier) answers where the request went; this answers why it was judged that way.
  • why: … — the classifier's own reason: either the sentence the LLM classifier gave, or the heuristic's scoring reasons. It answers "why this tier" where the route reason answers "which tier".

Two denominators, on purpose

The counters row counts requests; the Per turn line below it counts human turns.

Inside an agent loop one human message is re-sent on every tool step, and the classification is cached by bounded state and backend identity, so changed tool evidence can select a new tier. A per-request count is therefore weighted by how many steps a task happened to take — it mostly measures loop length, not the mix of work. A turn is identified by the last user message (its session plus its index in the request), so a five-step tool loop is one turn and a follow-up message is the next one.

Read the per-turn line to judge whether the difficulty mix is reasonable; read the per-request line to see how much traffic each tier actually served. When the two disagree sharply, the request denominator is being dominated by a few long loops.

The skipped note is the context guard at work: that tier was passed over because its model could not hold the request. The same data is on the stats endpoint:

curl -s localhost:3080/tier-router/api/stats | python3 -m json.tool

decisions is a bounded in-memory ring (20 entries) — it resets when the server restarts, and it records the router's own view, not a billed token count.

Long sessions and small-context models

A session grows, and a model that answered an early turn may no longer fit the current history — the classic failure is a local llama.cpp endpoint with a 131k window handed a 180k-token conversation:

400: request (180612 tokens) exceeds the available context size (131072 tokens)

contextGuard (on by default) handles that before the request is sent. For every candidate route the router resolves the model's context window (resolveModelInfo().context.contextWindow) and skips any route whose known window is smaller than the estimated request size, so the ladder falls through to a model that fits. Three properties matter:

  • Unknown windows never block. A provider that reports no window is always a candidate; the guard only ever drops a model whose limits it actually knows.
  • It never empties a chain. If every route would be skipped, the original chain is used unchanged — a routable request can never become "no route".
  • It is an estimate. The request size is approximated as ~1 token per CJK character and ~3.5 characters per token elsewhere (images add a fixed allowance), with 10% headroom. It is deliberately biased to over-estimate, because guessing low sends a request a model cannot hold.

Set contextGuard: false to route purely by difficulty and tier order. The model pickers also show each model's window (qwen3.8-27b-q5 · 131k ctx), so you can see which tiers can hold a long session when you configure them.

Configuration

Settings live in the tier-router namespace as flat fields. Edit them in the settings card or write them directly:

tier-router:
  enabled: true
  classifier: heuristic          # heuristic | llm | jev
  hardProvider: codex-local
  hardModel: gpt-6-astra
  hardEffort: ''                 # e.g. low / high / max; empty = unspecified
  normalProvider: codex-local
  normalModel: gpt-5.5
  normalEffort: ''
  easyProvider: gpudev
  easyModel: qwen3.8-27b-q5
  easyEffort: ''
  visionProvider: codex-local
  visionModel: gpt-6-astra
  visionEffort: ''
  visionMode: replace            # replace (structured evidence, default) | route (whole turn)
  visionCacheTtl: 3600           # seconds of vision-evidence cache; 0 disables
  visionFallbacks: []            # [{provider, model}] explicit vision fallbacks
  fallbackProvider: ''           # last resort; empty = the session default model
  fallbackModel: ''
  llmClassifierProvider: ''      # classifier: llm; empty = reuse the easy tier
  llmClassifierModel: ''
  # classifier: jev (TypeSafe System One). The key is normally pasted into the
  # settings card; TYPESAFE_API_KEY and ~/.typesafe/key are fallbacks.
  jevApiKey: ''                  # never echoed back to the settings card
  jevModel: jev-latest
  jevBaseUrl: https://api.typesafe.ai
  jevMinConfidence: 0.3          # below this Jev abstains and the heuristic decides (0 = off)
FieldDefaultMeaning
enabledtrueMaster switch; when off, requests go to the session default model.
classifierheuristicheuristic (built-in scoring), llm (a model generates a verdict) or jev (TypeSafe System One answers a typed choice).
hardScore3Heuristic only: score at or above which a request is hard. See the note below before changing it.
hardProvider / hardModel / hardEffortcodex-local / gpt-6-astra / ''Hardest tier.
normalProvider / normalModel / normalEffortcodex-local / gpt-5.5 / ''Everyday tier.
easyProvider / easyModel / easyEffortgpudev / qwen3.8-27b-q5 / ''Cheapest tier.
visionProvider / visionModel / visionEffortcodex-local / gpt-6-astra / ''Image tier.
visionModereplaceImage handling; see below.
visionCacheTtl3600Vision-evidence cache in seconds.
visionFallbacks[]Explicit vision fallbacks before the default model.
fallbackProvider / fallbackModel''Route used when no tier is configured; empty = session default.
llmClassifierProvider / llmClassifierModel''Classifier model for classifier: llm.
jevApiKey''TypeSafe key for classifier: jev. Falls back to TYPESAFE_API_KEY, then ~/.typesafe/key.
jevModeljev-latestModel alias for the Jev judgement.
jevBaseUrlhttps://api.typesafe.aiTypeSafe API base URL (/v1/systemone is appended).
jevMinConfidence0.3Below this confidence the Jev answer counts as an abstention and the heuristic decides. 0 disables the floor.
classifierTimeoutMs4000Budget for the semantic classifier (LLM or Jev). On timeout the heuristic decides immediately and the slow answer is cached for later requests.
visionTimeoutMs60000Budget for one vision-sidecar call, so a hung vision provider cannot stall the turn.
contextGuardtrueSkip routes whose known context window cannot hold the request.

The hardScore knob, and why it is not a fix

Measured on 213 real requests from one machine, the heuristic score distribution was degenerate:

ScoreRequests
61
21
110
0169 (79%)
-132

Nothing landed between 2 and 6, so hardScore values 2, 3, 4 and 5 behave identically. Only two settings do anything: 3 (the default — 1 request in 213 became hard) and 1 (12 did). 0 is safe for small talk because greetings score -1, and the schema clamps the value to 0..10 — a negative cut-off would classify every greeting as hard.

The knob is useful for choosing how aggressive to be, not for accuracy: 79% of requests sit in one bucket, so no threshold separates hard work from routine work. If routing quality is the goal, use a semantic classifier (classifier: llm or classifier: jev) instead.

How routing works

request ──► has images?
             ├─ yes ─► visionMode=replace: images → vision model → structured evidence text ─┐
             │         visionMode=route:   whole turn → vision tier                        │
             └─ no ───────────────────────────────────────────────────────────────────────┤
                                                                                            ▼
                                     difficulty classification (heuristic / LLM / Jev) ─► tier chain
                                              chosen tier → other tiers (nearest first) → default model

Difficulty signals are bilingual (English + Chinese): code-fence volume, number of file references, message length, and keywords such as architecture, refactor, concurrency, deadlock, memory leak, distributed, migration, performance, security…. A score of 3 or more is hard; a message containing a task verb is at least normal; greetings, translations and short explanations are easy.

Two behaviours worth knowing:

  • Image history counts. The request carries the whole conversation, so if any message contains an image the request is treated as a vision request — otherwise a text-only tier model would reject the history with "does not support image input".
  • visionMode. replace (default) keeps the vision model as an assistant only: it returns structured evidence, that text replaces the image, and the difficulty tiers produce the answer. route is the legacy behaviour: the whole turn goes to the vision tier.

When a model is unavailable

A tier that cannot serve a request is not the same as a tier that failed once. The router reads the harness failure code (HarnessError.code — the contract is explicit that callers route on the code and never parse the message) and reacts differently:

FailureCodesWhat the router does
The route cannot serveQUOTA, AUTH, INVALID_CREDENTIAL, MISSING_CREDENTIAL, NO_ADAPTERAdvance the chain and bench the route for routeCooldownMs (default 5 min).
Something went wrong onceTRANSPORT, TIMEOUT, SERVER, RATE_LIMIT, EMPTY_RESPONSE, unknownAdvance the chain, but bench only after routeFailureThreshold (default 2) consecutive failures inside the window.
About this request, not the routeCONTEXT_WINDOW_EXCEEDED, ABORTEDAdvance the chain and never bench: the next, smaller request may well fit.

A benched route is skipped outright — its failure is not paid again — and Settings → Tier Router lists it with the code, the message and the seconds until it is retried:

Benched routes (judged unavailable): codex-local/gpt-6-astra — QUOTA (usage limit reached), 274s retry in

Any success clears the route's streak, routeCooldownMs: 0 turns the bench off, and the bench is never allowed to empty the chain: if every route is benched the chain is used anyway, with the reason kept in the decision's skipped list (fail-open, like the context guard).

The streak rule exists because the code is not always informative. An exhausted account quota can surface as nothing but a connection timeout, because that is all the provider or its CLI ever said — measured on a real Codex CLI failure whose entire diagnostic was Reconnecting... 2/5 (request timed out). No amount of message parsing recovers a signal that was never emitted, so the router falls back to "the same route failed twice in a row".

The honest limit: this stops the bleeding from the second request onward. The first one still waits out whatever the adapter costs before it reports failure, and no failure code can arrive sooner than the failure does. Bounding that first request needs a timeout inside the adapter (for dsh-llm-codex, the timeoutMs field of its provider entry, default 10 minutes per codex exec).

HTTP API

The settings card talks to these host endpoints:

MethodPathPurpose
GET/tier-router/api/modelsModel catalog for every provider (image support, reasoning efforts) + current default model.
GET/tier-router/api/configResolved settings + defaults + writability.
POST/tier-router/api/configWrite one field ({field, value}); value: null resets it. Requires content-type: application/json and a same-origin Origin/Host pair, so a page you merely visit cannot rewrite the routing.
GET/tier-router/api/statsRoute counters (hard/normal/easy/vision/visionBridge/fallback/error/routeError) plus the per-turn counters in turns, the recent failure ring in errors, and the recent decision ring in decisions, and the routes currently benched in benched. error counts requests nothing could answer; routeError counts individual route failures, including the ones the fallback chain recovered.

The client half reads and writes through this API rather than the settings wire, because the host only exposes allow-listed namespaces to configuration clients.

Requirements

  • @deepseek-ai/dsh-llm ^0.1.5-rc.2 — adapter routing and contentHasImage
  • @deepseek-ai/dsh-settings ^0.1.5-rc.2 — installSection
  • @deepseek-ai/cordis ^4.0.2
  • @deepseek-ai/schemastery ^3.18.2
  • The client half contributes a settings.section entry and therefore loads only on platform: web.

Limitations and compatibility

Worth knowing before you rely on it:

  • The vision sidecar is cached, not free. Every request whose history contains an image calls (or reuses a cached call to) the vision model. Evidence is cached per attachment for visionCacheTtl seconds (default 1h); after that the same historical image is analysed again on the next request. Set visionCacheTtl: 0 to disable the cache.
  • An image anywhere in the history keeps the request image-bearing. That is deliberate (a text-only tier model would otherwise reject the whole history), and in replace mode the difficulty tiers still produce the answer — the vision model only supplies evidence. But a long session that once had a screenshot keeps paying for vision evidence until that turn leaves the conversation.
  • The heuristic is conservative. A hard verdict needs a score of 3 or more; two hard keywords alone land at normal. Retarget the tiers, or switch classifier to llm or jev.
  • This router only acts when you select it. It routes requests made through the smart model, so it is opt-in per session.
  • Do not run it together with a plugin that force-overrides the model on the agent/request waterfall (role-based routers that stamp planner/executor models after await next()). Such a plugin outranks the model selector, so Tier Router would never receive a request. Pick one.
  • Dependencies are peers. @deepseek-ai/dsh-llm, dsh-settings and cordis come from the DSH installation; the only bundled dependency is @deepseek-ai/schemastery.

Development

git clone https://github.com/zhangzhangco/dsh-tier-router
cd dsh-tier-router
npm install --legacy-peer-deps   # pulls the public @deepseek-ai/* peers
npm test                         # 202 cases, node:test, no test framework

npm test runs Node's built-in runner (node --test, auto-discovery). Note that node --test tests/ — the form upstream documented — does not work on Node 22.23: it treats the directory as a module path and exits with MODULE_NOT_FOUND.

To run the checkout in a profile instead of the published package:

dsh plugin --profile web remove dsh-tier-router
dsh plugin --profile web add link:/absolute/path/to/dsh-tier-router

Layout: index.js (bundle entry), lib/{schema,router,classifier,jev,vision,models-api}.js, lib/types/index.d.ts (hand-written TypeScript declarations), client/client.js (the settings card, a build-free window.__ModuleLoader__ bundle), cordis.patch.yml (the bundle layer), tests/.

Credits

This package is adapted from dsh-smart-router (MIT): the routing model, the difficulty classifier, the vision sidecar and the settings-card architecture originate there. This adaptation renames the plugin, retargets the tier defaults, fixes the client service injection, and drops the bundled free-vision provider seeding.

License

MIT.

State-aware classification

The default remains heuristic. Select llm and configure llmClassifierProvider/Model to use the generation-based classifier with task and step evidence, or select jev to let TypeSafe System One answer the same question as a typed choice (see below). Inputs omit private reasoning and plugin snapshots; actual downstream messages remain unchanged. agentStep counts current-turn tool calls, and failures require structured isError flags.

Two deterministic cases are not left to any classifier, including Jev: a short continuation (继续, continue, …) inherits the previous task's tier, and a repeated structured tool failure for the same call is at least hard — applied as a floor above the verdict, because the whole point of escalating is that the model in play already failed.

SHA-256 cache keys include the complete bounded input, backend/model and prompt. Concurrent identical work is shared; caller cancellation is isolated, late results stay keyed to their original state, and background work is capped at 30 seconds / 32 pending keys. Only bounded diagnostics are retained in memory.

A local option-scoring classifier (classifier: logits) was prototyped and then removed: its verdict flipped with candidate option order while still reporting near-maximum confidence, which makes the score unusable as a gate. See the decision record for the measurements and for what was not verified.

End-to-end task quality, latency and cost still need separate controlled runs; nothing in this repository establishes an accuracy claim for any of the three classifiers.

Jev (TypeSafe System One)

classifier: jev asks TypeSafe's hosted judgement model one choice question — which tier handles the next step? — and gets back the chosen tier, a probability per option and a confidence. Nothing is generated, so there is no reply format to misparse; the router's policy stays in code.

Setup, in order of precedence:

  1. paste the key into Settings → Tier Router → Jev (jevApiKey), or
  2. export TYPESAFE_API_KEY, or
  3. rely on ~/.typesafe/key if you already have the TypeSafe SDK installed.

The key is stored in this machine's settings and is never echoed back to the settings card: the config API reports only whether a usable key exists and where it came from. Because the card cannot read the stored key, the input is always blank — leave it blank to keep the current key, or paste a new one to replace it.

What was measured on the development machine (Apple M4, macOS, to api.typesafe.ai): 0.73–0.82s per round trip end to end (connect ~0.2s) for a three-option question over a small state object, against the classifierTimeoutMs budget of 4000ms. That is the cost you pay once per changed turn — the decision is cached by bounded evidence plus question wording and model, so the tool steps of one loop do not each pay it.

Failure handling is deliberately boring: a missing key, a rejected key (401/403), a rate limit (429/529), a timeout, a transport error or an unusable answer all resolve to a named jev-* classifier in the decision record and hand the decision back to the heuristic. A refused key is additionally not retried for 60 seconds, so a wrong key does not cost a round trip on every turn.

jevMinConfidence is the abstention floor. A uniform distribution over three options has a confidence near 0, so the default 0.3 rejects exactly the judgements that carry no real signal rather than letting a coin flip pick a tier. Set it to 0 to trust every answer, or raise it if you would rather have the heuristic decide more often.

The rubric itself — the instructions and the per-option what / not_for / examples — lives in lib/jev.js. Its wording is part of the cache identity, so editing it can never serve a decision made under the old phrasing.

Comments

Loading…

Similar plugins

dsh-smart-router

by rouyiemei

Automatic model routing for DeepSeek Harness: three difficulty tiers (hard/normal/easy) plus vision routing, picking models you already configured under Settings ? Models. ???? + ?????????

Manifest valid

★ 0

MIT

JavaScript

Aug 17, 2026

dsh plugin --profile web add dsh-smart-router

Automatic model-tier routing for DeepSeek Harness: one user instruction enters, one tier decision comes out — complex intent is planned on the strong tier and implemented on the cheap tier, simple int

Models & ProvidersManifest valid

★ 1

↓ 932/wk

Apache-2.0

TypeScript

dsh plugin --profile web add dsh-autotier

by llmpolska

Tiered model routing for DeepSeek Harness: think/build model tiers, vision delegation (the vision model only describes images; the working model acts), image generation, and a high-impact guard.

Models & ProvidersTools & CapabilitiesManifest valid

★ 1

↓ 161/wk

MIT

JavaScript

Aug 16, 2026

dsh plugin --profile web add oh-my-dsh

by BruceLanLan

Two-tier model routing: a strong tier plans, advises and reviews while a cheap tier implements, with plan-mode-aware auto routing, a high-impact escalation guard, failure auto-escalation, and subagent

Models & ProvidersWorkflow & AutomationTerminal & ClientsManifest valid

★ 6

↓ 367/wk

MIT

JavaScript

Sep 3, 2026

dsh plugin --profile web add dsh-tier-router

Unified model routing for DeepSeek Harness: one logical ModelID over multiple providers with first-token failover and cooldown, health-aware candidate ranking, three tiers (tier1/tier2/tier3) auto-sel

Models & ProvidersManifest valid

★ 0

↓ 384/wk

dsh plugin --profile web add @welsione/dsh-model-router

by ringoage

The visual, manual subagent-model picker for DeepSeek Harness — predictable per-session routing with global coverage of every in-process subagent path, complementary to auto-routing plugins.

Models & ProvidersWorkflow & AutomationManifest valid

★ 0

MIT

JavaScript

Aug 16, 2026

dsh plugin --profile web add dsh-subagent-model-picker