DSH Plugins Marketplace

DSH Plugins

Plugins

/

Development & Infrastructure

/

dsh-voice-transcribe

J

dsh-voice-transcribe

Manifest valid

dsh plugin: local speech/video transcription — QQ/WeChat-style SILK voice is decoded first, then converted to text using faster-whisper, with no API cost throughout · Local speech-to-text for dsh: SILK decoding + faster-whisper, no API cost.

hasBundlePatch

dsh-voice-transcribe

English | 简体中文

给 DeepSeek Harness(dsh)用的本地语音/视频转写插件: 让 Agent「听得见」——把一个音频或视频文件交给它,它把里面的话转成文字。

引擎是 SenseVoice(sherpa-onnx 跑 int8 onnx)。全程本地跑,不花 API 钱, 约 0.1 秒 / 条,模型 239 MB。

同类里更成熟的选择:桌面麦克风那类(比如 likhonmain/voice-input)比这份成熟得多—— 你要的是「对着电脑说话、转成字」,去装那些。 我们这份专做别人在 QQ / 微信里发过来的语音:SILK 解码、QQ 文件头多出来的那个字节、本地 SenseVoice 转写,一条龙。 还有一条事实要说清:官方 QQ 机器人路线不需要这一套——平台自带识别文字(asr_refer_text);只有个人号路线(OneBot:NapCat / Lagrange / SnowLuma)才用得着它。


为什么需要它

它真正解决的问题,不是「转写」。

调模型谁都会。麻烦的是前面那一步,尤其是中文 IM 生态里:

坑 1:QQ / 微信发出来的语音是 SILK,不是 amr

QQ 里真人语音的文件头是 #!SILK_V3,但文件名常写着 .amr。 你以为喂给 ffmpeg 就行——不行。

坑 2:大多数 ffmpeg 构建根本没有 SILK 解码器

$ ffmpeg -decoders | grep silk
# (空)
$ ffmpeg -i voice.amr out.wav
# Invalid data found when processing input

看着像文件坏了,其实是没人认这个格式。于是你开始怀疑是不是文件截断了、 是不是要用 QQ 的私有库——都不是。

坑 3:QQ 存下来的文件常常在最前面多一个字节

用十六进制看:

00000000: 0223 2153 494c 4b5f        .#!SILK_

那个 02 是 QQ 自己塞的(03 也见过)。不剥掉它,连专门的 SILK 解码库都会拒绝你的文件。

这三件事都在 py/silk.py 里处理掉了——用 pilk(纯 Python,不需要编译任何东西)解成 PCM, 自己封标准 WAV。顺手把一个 .amr 丢进去也能用,is_silk() 会告诉你它到底是什么。

更完整的"踩坑笔记"(适用范围、怎么给 SILK 写一个不作假的自检)在 docs/silk-notes.md。

那用哪个引擎?

whisper 谁都会调,我们一开始也是它。同一条 1.7 秒的真实群语音: whisper-medium 试了 5 组参数(beam 1/5 × 词表提示 / 口语 / 无提示),全部听岔; SenseVoice 一次就对,和官方转写一字不差。速度也不是一个量级:约 0.1 秒 / 条(whisper 5-8 秒), 模型 239 MB int8(whisper medium 约 1.5 GB)。

所以这个插件里只有 SenseVoice 一条路:whisper 的开关、参数、依赖都删干净了, 不留半截开关让人以为还能切回去。数字都在下面「测试(实测数据)」里。

安装

1. 装插件

dsh plugin --profile web add github:JackZo400/dsh-voice-transcribe

2. 装 Python 那半边(装在 pythonPath 指的那个解释器里)

pip install -r py/requirements.txt      # sherpa-onnx + numpy + pilk

镜像提醒:清华源(pypi.tuna.tsinghua.edu.cn)里没有 sherpa-onnx 这个包,会直接报 No matching distribution found(实测过)。装不上就换阿里源或官方源: pip install -i https://mirrors.aliyun.com/pypi/simple -r py/requirements.txt

3. 系统里要有 ffmpeg(非 SILK 的音频/视频靠它转码)

4. 下 SenseVoice 模型(约 163 MB 的压缩包,解开是 239 MB 的 int8 模型)

curl -LO https://github.com/k2-fsa/sherpa-onnx/releases/download/asr-models/sherpa-onnx-sense-voice-zh-en-ja-ko-yue-int8-2024-07-17.tar.bz2
tar xjf sherpa-onnx-sense-voice-zh-en-ja-ko-yue-int8-2024-07-17.tar.bz2

解出来的 sherpa-onnx-sense-voice-zh-en-ja-ko-yue-int8-2024-07-17/ 就是模型目录, 里面要有 model.int8.onnx 和 tokens.txt。放哪都行,但要写进配置的 modelDir (或环境变量 AILIN_ASR_SV_DIR)——代码里没有任何写死的路径。 这个模型认中文(普通话)/ 英文 / 日文 / 韩文 / 粤语五种语言, 想换别的版本看 sherpa-onnx 的模型发布页。

配置

- insert:
    - id: voice-transcribe
      name: dsh-voice-transcribe
      config:
        pythonPath: python3        # 装了上面那些包的解释器
        # scriptPath: ''           # 留空 = 用包内自带的 py/transcribe.py
        modelDir: /path/to/sherpa-onnx-sense-voice-zh-en-ja-ko-yue-int8-2024-07-17
        maxSeconds: 120            # 只转前 N 秒
        language: auto             # auto / zh / en / ja / ko / yue
        fix: ''                    # 专有名词纠错表,见下
        maxFiles: 4                # 一次最多转几个(模型只加载一次)

modelDir(或环境变量 AILIN_ASR_SV_DIR):模型目录在哪由你说了算,代码不猜路径。 没配的话转写会明确报「没给 SenseVoice 模型目录」,不会静默失败。

fix(或环境变量 AILIN_ASR_FIX):专有名词纠错表。SenseVoice 没有词表接口 (whisper 的 initial_prompt 那种),名字听岔了只能在结果上纠一道。格式 错=对,错=对:

fix: '小张=张三,星海=星海项目'

表是从左往右替换的,短词会吃掉长词的前缀——长的写前面。 默认是空表:谁的名字谁自己填。

用法

在 dsh 里(Agent 自己调)

装好之后 Agent 多一个工具 transcribe_media:

把 /tmp/voice.amr 转成文字

返回每个文件一条结果:file / ok / text / language / duration。

一次也可以给多个文件——模型只加载一次,批量比一个个转快得多。

别的插件想用,可以调服务:

const svc = ctx.get('voiceTranscribe')
const { results } = await svc.transcribe(['/tmp/voice.amr'])

单独用(不需要 dsh)

只要 SILK 解码:

python py/silk.py --check voice.amr        # 先看看是不是 SILK
python py/silk.py voice.amr --out-dir out/ # → out/voice.wav(16k 单声道)

要转写:

python py/transcribe.py zh.wav --language zh --model-dir /path/to/sherpa-onnx-sense-voice-zh-en-ja-ko-yue-int8-2024-07-17
# {"file": "zh.wav", "dur": 5.59, "ok": true, "text": "开饭时间早上9点至下午5点。", "engine": "sensevoice", "lang": "zh", "sec": 1.0, "load_sec": 0.7, "total_sec": 1.0}

(上面这行输出是模型自带测试音频 test_wavs/zh.wav 的真实结果,你可以自己复现。)

transcribe.py 的输出是一行 JSON,方便被任何程序调用(dsh 插件就是这么接的): 出错时退出码仍是 0,错误放在 JSON 的 error 字段里。

测试

python py/test_silk.py          # SILK 识别的逻辑(不需要真语音样本,秒级)
python py/test_transcribe.py    # 转写那条路:参数/纠错/懒加载;没模型会自动跳过
python py/test_silk_real.py     # 用 pilk 现造真 SILK,验解码和 QQ 那个 0x02 前缀
node test/plugin-selftest.mjs   # 插件接线(用假转写脚本,不需要模型)

py/test_silk.py 覆盖的就是那三个坑:干净头、0x02/0x03 前缀、空文件、 mp3 头、短文件不崩、不是 SILK 时抛错。

py/test_transcribe.py 不需要模型也不需要联网:先验参数解析和纠错表, 再拿桩替掉 sherpa_onnx 走一遍完整流程(顺便证明它真的是用到才 import), 最后装了 sherpa-onnx 且配了模型目录才会再跑一遍真模型。 没有模型、没装包的环境下它只打 SKIP 并说明原因——不会假装通过。

py/test_silk_real.py 补的是前两个够不着的那一段:仓库里不带音频,它就用 pilk 的 encoder 现造一个真 SILK(音源优先取 SenseVoice 模型包自带、可公开分发的测试音频,没有就自己合成一段音调), 再验 QQ 那个「前面多一个 0x02」的坑——人为加一个字节后解出来的 PCM 和不加那次逐字节一致; 装了 sherpa-onnx 且配了模型目录,它还会把解出来的声音真转一遍。缺依赖时同样只打 SKIP 并说清缺什么,退出码仍是 0。

想看真效果,拿手上任意一条 QQ/微信语音跑 python py/silk.py 你的文件。

测试(实测数据)

本机(14 核 CPU、int8)真实语音样本:

项目数字
模型体积239 MB(int8 onnx;whisper medium 约 1.5 GB)
模型加载约 0.9 秒(一个进程只加载一次;whisper medium 约 1.6 秒)
同一条 1.7 秒真实群语音约 0.1 秒出字,一次就对(whisper-medium 试了 5 组参数全听岔)
SILK 解码约 0.2 秒 / 条
3.6 秒真实语音出字:「那样人太刷屏了 我直接给踢了不是说了吗」

那条 1.7 秒的语音,SenseVoice 和 QQ 官方转写一字不差;whisper-medium 的 5 组参数一组都没对。 零 API 花费,费的是 CPU。CPU 越强越快。

已知局限

  • 只吃本地文件。URL 请先自己下载——插件不替你做网络请求(少一个 SSRF 面)。
  • 靠子进程:模型崩了、超时了都杀得掉,但它不共享内存,每次调用有一次 Python 启动开销。
  • 视频只转音轨,不抽帧——画面理解是另一件事。
  • 纯静音 / 纯音乐不保证是空串:SenseVoice 的幻觉比 whisper 收敛得多,但实测给一段纯数字静音, 它偶尔还是会冒一两个没意义的字。空的就如实是空的,但别把「有字」当成一定有话。
  • SenseVoice 没有词表接口:专有名词只能靠 fix 那张表兜,得自己攒。
  • 长音频靠切片:超过 28 秒会在最安静的地方切片再拼;切点挑得再小心也偶尔会把词切开。
  • 只认 5 种语言:中文(普通话)/ 英文 / 日文 / 韩文 / 粤语,别的语言请另找模型。

License

MIT © 2026 JackZo400


English

→ Full English README: README.en.md

Comments

Loading…

From the same category

awesome-dsh-plugin

by awesome-dsh-plugin

A curated list of plugins for DeepSeek Harness (dsh) · DeepSeek Harness 插件精选列表

Development & Infrastructure

★ 17.5k

CC0-1.0

Python

Oct 1, 2026

Index only — not installable

by zhu1090093659

DeepSeek Harness (DSH) Web Plugin Aggregation Ecosystem · Everything is a plugin, distributed via the Creative Workshop

Tools & CapabilitiesUI & ExperienceDevelopment & InfrastructureTerminal & ClientsModels & ProvidersManifest valid

★ 8.3k

↓ 172/wk

Apache-2.0

TypeScript

Oct 1, 2026

dsh plugin --profile web add dsh-web

by yjh051108

dsh-routing-suite — injector + router-standard kit: install the runtime injector first, then the task-aware reasoning-mode router preset (measured P1-P23).

Development & InfrastructureModels & ProvidersManifest valid

★ 7k

MIT

JavaScript

Sep 18, 2026

dsh plugin --profile web add @dsh-external/dsh-super-injector

by strukto-ai

The World's First Virtual Terminal for AI Agents

Tools & CapabilitiesDevelopment & InfrastructureManifest valid

★ 3.7k

↓ 264/wk

Apache-2.0

TypeScript

Oct 1, 2026

dsh plugin --profile agent add @struktoai/mirage-dsh

by xmanrui

通过扫码或机器人凭据把IM机器人接入DeepSeek Harness(支持飞书、微信、钉钉、企业微信、QQ、Slack、Telegram、Discord和WhatsApp)。 Connect IM bots to DeepSeek Harness via QR code or credentials (9 channels).

Tools & CapabilitiesNotifications & IntegrationsDevelopment & InfrastructureManifest valid

★ 1.6k

↓ 17.5k/wk

MIT

JavaScript

Oct 1, 2026

dsh plugin --profile web add @xmanrui/dsh-im

by hyhmrright

AI code reviews grounded in 12 classic engineering books — decay risk diagnostics with book citations, severity labels, and 6 analysis modes including full-sweep auto-fix

Tools & CapabilitiesDevelopment & Infrastructure

★ 1.5k

MIT

JavaScript

Sep 28, 2026

Index only — not installable