dsh-session-slm-router
Manifest validShadow-mode SLM router for DeepSeek Harness: per-turn weak/strong prediction via vertical-small-model CLI, writes shadow JSONL log
dsh-session-slm-router
Shadow-mode SLM 路由插件 for DeepSeek Harness。每轮通过 vertical-small-model CLI 预测用户发言的强弱分类(weak/strong),记录影子日志 JSONL,不改变会话活跃模型。
⚠️ S1 阶段:默认
shadow模式,只观察不换模。weak-only模式(S3 灰度)可执行降档换模,仅放行 weak→strong 降档,不改变用户原模型持久化。
安装
# 通过 dsh plugin 安装(推荐,自动处理 bundle 注册)
dsh plugin add dsh-session-slm-router
# 或手动添加到 profile package.json 的 dependencies + dsh.profile.bundles
配置
模式选择
| mode | 行为 |
|---|---|
shadow(默认) | 只预测 + 写影子日志,不换模 |
weak-only(S3 灰度) | 真正换模,但只放行 switch_to_weak(降档);switch_to_strong 记录但不放行 |
off | 关闭插件 |
weak-only 绑定时做 reasoning effort 能力兼容预检(宿主对不兼容组合硬拒):目标无该能力 → 当轮剥离请求参数后绑定;值不在目标支持列表 → 顺延次选;全部不兼容 → 不绑定(不降档)。详见「数据」节说明。
同意门(
downgradeConfirm):bind 前可经宿主审批通道弹「是否降档」确认卡(浏览器dsh-client-ui-approval渲染,标题=reason);always每次 bind 都问,smart仅高意图情景(会话首轮/强模型≠默认)问,never静默。拒绝/超时(30s)/无 UI → 当轮不降档(安全方向),结果落downgrade_confirm影子字段。
完整配置示例
# 仅 cordis 行内 config(bundle 的 cordis.patch.yml,或宿主 profile 的 cordis.patch.yml 覆盖行)
# ⚠️ settings.yaml 的插件命名空间块不会注入插件 config(2026-09-03 实测)
session-slm-router:
mode: shadow # shadow | weak-only | off
weakOnlyMainOnly: true # weak-only 是否只对主会话换模
predictCmd: "python3 <path-to>/vertical-small-model/scripts/route_predict.py"
predictModel: "<path-to>/vertical-small-model/data/eval/routing-v0/model-r1.json"
timeoutMs: 250
logPath: "slm-shadow/session-slm-shadow.jsonl"
slotCachePath: "slm-shadow/slot-order-cache.json" # 可选:槽位缓存路径(相对 ~/.dsh);缺省=与 logPath 同目录
tierRules:
maxTokensWeakMax: 32768 # L2 元数据降档阈值
autoWeak: true # 自动收录:元数据确认的弱模型进候选池(显式 strongSlots 否决)
freeFirst: false # free 优先排序(默认关:free 供应商普遍限额)
downgradeConfirm: never # 同意门:never(默认,静默自动降档)| always(每次 bind 弹「是否降档」确认卡)| smart(仅高意图情景弹)
weakSlots:
- { provider: <provider>, model: <model> }
strongSlots:
- { provider: <provider>, model: <model> }
predictCmd/predictModel:必填,指向 vertical-small-model 项目的 CLI 脚本与模型文件weakSlots/strongSlots:运行时模型槽位候选表,按你想使用的 provider/model 填写tierRules.autoWeak(默认开):新供应商/模型(如其他 agent 直接改 settings.yaml 加入)自动收录——元数据确认(maxTokens ≤ 阈值)即进降档候选池,显式strongSlots一票否决;free 关键字不参与强弱推断freeFirst(默认关):free 模型/供应商普遍限额,free 优先仅作为显式开启的排序偏好slotCachePath(可选,2026-09-09 缓存隔离):槽位健康缓存文件路径(相对~/.dsh,与logPath同约定)。缺省 =dirname(logPath)/slot-order-cache.json(与日志同目录,即默认~/.dsh/slm-shadow/)。测试/隔离场景显式指向独立目录,避免把测试缓存写进生产目录(scripts/mount.mjs即经此键隔离)downgradeConfirm(默认never,2026-09-09 同意门):weak-only bind 前是否经宿主审批通道(浏览器dsh-client-ui-approval渲染确认卡,标题「是否降档」)征得用户同意。always=每次 bind 都弹;smart=仅高意图情景(会话首轮 / 当前强模型≠默认选择)弹;never=静默自动降档(现状)。超时 30s;拒绝/超时/无 UI 一律不降档(fail-closed,安全方向),结果落影子日志downgrade_confirm字段- 配置只读
cordis.patch.yml行内 config:bundle 自带注册行 + 宿主 profilecordis.patch.yml的- id: session-slm-router+config:覆盖行(本机文件,可含本机路径)
槽位健康缓存(weak-only 换模顺序)
weakSlots 的配置顺序 ≠ 生效顺序。插件维护一个持久化健康缓存:
- 缓存文件:
~/.dsh/slm-shadow/slot-order-cache.json(=slotCachePath,缺省为与logPath同目录;删除即重建;仅持久化显式 weakSlots,auto 收录槽不落盘、重启重推) - 生效顺序:可用槽(
freeFirst: true时 free 槽优先,默认按配置序)→ 瞬态未知槽(原配置序)→ 死亡槽(沉底);tierRules.autoWeak: true时显式 weakSlots 之后自动追加元数据确认的弱模型 - 死亡判定:provider 已注销 / 模型不在该 provider 已配置清单(catalog)/ 模型实测不可用(404 / NO_ADAPTER 等)→ 沉底
- 刷新时机:启动(拓扑未就绪时 2s/4s/6s 有限重试)、
llm/adapters-updated、agent/session-start、24h TTL 过期;绑定路径发现元数据缺失时惰性补收集(重挂载兜底)
数据
影子日志写在 ~/.dsh/slm-shadow/session-slm-shadow.jsonl(一行一条 JSON),字段:
| 字段 | 说明 |
|---|---|
v | schema 版本(当前 1) |
ts | 时间戳 |
session_id | 会话 ID |
turn_seq | 轮次序号 |
utterance_hash | 用户发言 sha256 前 16 位(脱敏) |
utterance_preview | 前 80 字(脱敏) |
suggested_tier | CLI 预测强弱结果 |
confidence | 预测置信度 |
actual_tier | 当前实际模型档位 |
actual_tier_cause | 实际档裁决来源:registry-weak / registry-strong / metadata / heuristic / abstain-unknown-model |
target_source | 降档目标来源:registry(显式 weakSlots)/ auto(自动收录);未绑定为 null |
switch | 决策结果:stay / switch_to_weak / switch_to_strong |
bound | weak-only 是否实际换模(shadow 恒 false) |
downgrade_confirm | 同意门结果(2026-09-09):confirmed 用户同意 / denied 拒绝 / cancelled 超时或取消 / unavailable 无 UI 或无应答器 / not-required 未开启同意门;非 bind 轮次为 null |
predict_ok | 预测是否成功 |
error | 预测失败原因摘要 |
kind | 仅 bind-failure 事件有(bind-failure):降档目标执行层失败回放(reasoning effort 不兼容,宿主 prepareCall 硬拒);轮次事件无此字段。本事件不含 utterance 字段(数据源是执行层错误,无用户话语) |
reasoning_effort | 仅 bind-failure 事件有:请求携带、目标模型不支持的 reasoning effort 值 |
2026-09-09 起,weak-only bind 前做 reasoning effort 能力兼容预检(通用机制,按宿主
resolveModelInfo().reasoning声明能力判定,非模型白名单):目标确认无该能力 → 当轮剥离请求中的reasoningEffort再绑定;目标支持 reasoning 但不含该值 → 顺延次选;候选元数据未知 → 不猜顺延;全部不兼容 → 不绑定(安全退化=不降档)。目标执行仍失败时经agent/error落kind=bind-failure事件(兜底记录,不淘汰槽位——模型没死,只是参数不兼容)。
开发
git clone https://github.com/NinjaSln-labs/dsh-session-slm-router.git
cd dsh-session-slm-router
npm install --legacy-peer-deps # peer 由宿主 dsh 提供
npm run build # tsc + build.mjs 探测包装 → lib/
npm test # 验证链单源(scripts/verify.mjs):build → typecheck → unit → mount
npm run mount # 单跑挂载冒烟测试(验证 apply 运行 + 影子管线端到端;已含于验证链)
测试
验证链单源 scripts/verify.mjs(npm test / CI / publish 三处同调),测试使用 Node 内置 test runner,零外部依赖,不触网:
node --test tests/ # 112 用例:switch 判定表 / health 标记 / would_bind 语义 / 五级证据链 tierOf(弱档词整词匹配)/ autoWeak 收录 / 缓存隔离 / reasoning effort 兼容 / 同意门 / 超时降级
依赖
- 运行时:
@deepseek-ai/cordis(dsh 宿主) - 外部:
vertical-small-modelCLI(预测器,需独立部署) - 开发:TypeScript + Node 内置 test runner
贡献
License
MIT © 2026 ninjasln
Comments
Loading…
Similar plugins
by icyaaaww
Deterministic per-turn adaptive model routing for DeepSeek Harness
★ 0
MIT
JavaScript
Aug 24, 2026
dsh plugin --profile web add dsh-adaptive-model-routerby qinyu765
Provider discovery, matching, health checks, and pre-output failover for DeepSeek Harness
★ 0
MIT
TypeScript
Aug 14, 2026
dsh plugin --profile web add dsh-llm-auto-routeby jhuanxx44
LLM debug console for DeepSeek Harness: captures the full content of every model call at the llm/stream waterfall (options, stream chunks, usage, wire protocol) and shows it in a DevTools-style docked
★ 1
↓ 289/wk
MIT
TypeScript
Sep 16, 2026
dsh plugin --profile web add dsh-sseyeby ardli-firman
Searchable model selector for DeepSeek Harness — search models by name instead of scrolling
★ 0
JavaScript
Aug 22, 2026
dsh plugin --profile web add @ardli-firman/dsh-model-searchby striveh
Local request and response inspector for session-associated DeepSeek Harness LLM calls
★ 0
MIT
TypeScript
Aug 31, 2026
dsh plugin --profile web add dsh-llm-call-inspectorby zhangzhangco
Automatic tier-based model routing for DeepSeek Harness (dsh): a virtual `smart` model classifies every request by difficulty (hard / normal / easy) and by vision need, then delegates it to the models you already configured.
★ 0
NOASSERTION
JavaScript
Sep 22, 2026
dsh plugin --profile web add dsh-tier-router