DSH Plugins Marketplace

DSH Plugins

Plugins

/

dsh-voice

w

dsh-voice

Manifest valid

plugin of voice input for DeepSeek Harness

UI (client)hasBundlePatch

dsh-voice

Voice input for DeepSeek Harness (dsh).

A single dual-face plugin that brings local, private speech-to-text to both the Web UI and the TUI:

  • Host half — a ctx.stt service backed by an ONNX Whisper model (@huggingface/transformers + onnxruntime-node), plus the HTTP routes the browser mic button posts to.
  • Browser half — a microphone button in the composer (exports["./client"]).

No external binaries are required: the ONNX runtime ships as a prebuilt native addon per platform, and the model is a quantized Whisper ONNX downloaded from the Hugging Face Hub on first use. Microphone recording in the TUI still needs a system recorder (ffmpeg / sox / arecord); web voice records in the browser and needs nothing extra.

Install

dsh-voice is a bundle you add to a dsh profile. From npm:

dsh plugin --profile web add dsh-voice

Or straight from this repository (pnpm runs the prepare build on install):

dsh plugin --profile web add github:wencharmwang/dsh-voice

pnpm ≥10 refuses to run a git dependency's prepare script until it is allowlisted. If the first add fails, copy the printed package key into the profile's pnpm-workspace.yaml under allowBuilds and re-run the add.

The bundle declares its own cordis.patch.yml, so dsh plugin add both installs the package and activates the voice row. For a manual profile you can declare the row yourself in ~/.dsh/profiles/web/cordis.patch.yml:

- insert:
    - id: voice
      name: dsh-voice

Configuration

All fields are optional. Configure them either on the bundle row, or — once the voice row exists — by patching it by id in the profile's cordis.patch.yml:

- id: voice
  config:
    model: onnx-community/whisper-medium  # a Hugging Face Whisper ONNX id | a local dir
    language: auto                       # 'auto' | 'zh' | 'en' | ...
    dtype: q8                            # 'q8' | 'fp32' | 'q4'
    dir: ''                              # model cache + recordings dir; empty = ~/.dsh/voice

The effective model is resolved in two layers:

  1. Schema default — onnx-community/whisper-medium (q8, downloaded on first use).
  2. cordis.patch.yml — the plugin's model config overrides the default per deployment.

Switching models downloads the new one on first use after the switch; the superseded pipeline is disposed once no transcription is using it.

Service API (ctx.stt)

interface SttService {
  /** Readiness + effective model id, without loading or downloading anything. */
  status(): { ready: boolean; preloading: boolean; model: string }
  /** Ensure the effective model is downloaded and loaded. */
  ensureModel(onProgress?: (done: number, total?: number) => void): Promise<string>
  /** Transcribe audio input to text. */
  transcribe(
    input: string | Buffer | Float32Array,
    options?: { language?: string; signal?: AbortSignal; format?: string },
  ): Promise<string>
  /** Record from the microphone to a WAV file (TUI only). */
  startRecording(): { readonly path: string; stop(): Promise<string>; cancel(): Promise<void> }
}

transcribe accepts a Float32Array (16 kHz mono), a Buffer of PCM/WAV bytes, or a path to a WAV file. Use format to hint the byte layout: pcm16 (default), f32, or wav.

HTTP routes (web)

RouteMethodDescription
/voice/transcribePOSTBuffer the uploaded audio and return the transcript as { "text": "…" }. Optional ?lang=<iso> pins the language.
/voice/statusGET{ "ready": true } when the model is loaded.

The browser mic button records, decodes to 16 kHz mono PCM in the browser, and POSTs it as audio/l16;rate=16000.

Models

idapprox. sizenotes
onnx-community/whisper-small~250 MBlighter option
onnx-community/whisper-medium~0.8 GBq8, the default
onnx-community/whisper-large-v3~1.6 GB

Models are cached under ~/.dsh/voice/models after the first download.

Developing

This plugin is a self-contained bundle — it builds independently with pnpm install && pnpm run build (no monorepo checkout required). The prepare script runs the same build on git/tarball installs. See PUBLISHING.md for how to publish it to the dsh-plugin community.

License

MIT

Comments

Loading…

Similar plugins

dsh-voice-input

by forrestahha

Voice-to-text input plugin for the DeepSeek Harness Web UI

Terminal & ClientsDevelopment & InfrastructureManifest valid

★ 2

↓ 57/wk

MIT

TypeScript

Aug 14, 2026

dsh plugin --profile web add dsh-voice-input

by zemanzhang809

A speech-to-text (voice input) plugin for DeepSeek Harness.

Development & InfrastructureManifest valid

★ 1

MIT

TypeScript

Sep 18, 2026

dsh plugin --profile web add dsh-stt-plugin

by yafangwang9

Voice input plugin for DeepSeek Harness

Sessions & MessagesDevelopment & InfrastructureManifest valid

★ 3

↓ 41/wk

MIT

JavaScript

Aug 24, 2026

dsh plugin --profile web add dsh-voice-plugin

by zhuiyueya

Voice for DeepSeek Harness(dsh) — speech-to-text input + read-aloud TTS for text-only DeepSeek, zero API key.

Manifest valid

★ 4

↓ 436/wk

MIT

JavaScript

Aug 15, 2026

dsh plugin --profile web add dsh-voice

by 3274375092

Voice input plugin for DeepSeek Harness: mic → local/browser speech recognition → text submitted as a normal chat message. Input-only and preset-agnostic.

UI & ExperienceManifest valid

★ 8

↓ 158/wk

MIT

TypeScript

Oct 1, 2026

dsh plugin --profile web add @nn12138/dsh-voice

by allmodels-io

Real-time voice input and spoken answer summaries for DeepSeek Harness.

UI & ExperienceManifest valid

★ 3

↓ 123/wk

MIT

TypeScript

Sep 22, 2026

dsh plugin --profile web add @allmodels/dsh-speech