DSH Plugins Marketplace

DSH Plugins

Plugins

/

dsh-voice-input

m

dsh-voice-input

Manifest valid

Offline local speech-to-text for the DeepSeek Harness composer: a microphone button that records, transcribes with faster-whisper on your own machine, and inserts the text into the chat draft.

UI (client)hasBundlePatch

dsh-voice-input

Offline local speech-to-text for the DeepSeek Harness composer.

A microphone button appears in the chat composer. Click it, speak, click again, and the recognized text is appended to the draft — so you can fix a misheard word and press Enter yourself. Nothing is uploaded: audio is recorded in your browser, posted to the harness process on your own machine, and decoded there.

The model loads once and stays warm, so later recordings cost only the decode. Measured on this machine with base.en on a Ryzen 7 7800X3D, warm: ~0.4 s per recording on the CPU engine, ~0.1 s on the GPU engine. Before that change every recording reloaded the model, which took ~10 s.

Text appears in the draft while the audio is still being decoded, so a long instruction can be watched rather than waited for.

Engines

CPU (default)GPU (whisper.cpp)
decoderfaster-whisper / CTranslate2whisper.cpp
acceleratorCUDA if an NVIDIA card is presentVulkan (AMD, Intel, NVIDIA)
warm latency, base.en~0.4 s~0.1 s
first run~0.7 s model load1.7–3.3 s once, while the driver compiles Vulkan pipelines
in-progress textprogressive, while decodingone update when the transcript is ready
GPU-resident statenoneone model copy (~418 MB) per running engine

The CPU engine is the default: it is already sub-second warm, its text appears progressively, and it holds no GPU state that could accumulate. The bundled whisper.cpp engine is there because CTranslate2 — and therefore faster-whisper — accelerates only through CUDA, which leaves AMD cards on Windows unreachable without it. Set backend: auto (or ggml to require it) in the row to use it.

Requirements

  • A DSH installation with the Web GUI (the web profile).
  • Python 3.9+ on PATH for the one-time engine setup (python on Windows, python3 elsewhere).
  • uv is optional; the setup uses it when present and falls back to pip.
  • A Chromium-based browser or Firefox with microphone recording support.
  • For the GPU engine only: a Vulkan-capable GPU and a Windows whisper.cpp Vulkan build (bundled here). Its GGML model downloads on first use (~142 MB).

Install

# 1. Install the bundle into your profile.
dsh plugin --profile web add github:moazzamak/dsh-voice-input

# 2. Build the local engine once. It lives inside the installed package, which
#    creates a venv there and pre-downloads the default model.
python "$HOME/.dsh/profiles/web/node_modules/dsh-voice-input/python/setup.py"

On Windows (PowerShell):

dsh plugin --profile web add github:moazzamak/dsh-voice-input
python "$env:USERPROFILE\.dsh\profiles\web\node_modules\dsh-voice-input\python\setup.py"

Then restart DSH. Note that dsh web is a fixed alias for the web profile — dsh web --profile other does not exist, and dsh plugin must target the same profile you actually boot.

Alternatives to a git install

A git install fetches sources, and pnpm ≥10 refuses to run a package's prepare script until you allow it. This package ships plain JavaScript with no build step, so there is nothing to run and nothing to allow.

  • Tarball: pnpm pack in this directory, then dsh plugin --profile web add ./dsh-voice-input-0.1.0.tgz.
  • Local checkout: dsh plugin --profile web add /path/to/dsh-voice-input.
  • npm: dsh plugin --profile web add dsh-voice-input (if published).

Use

In the Web GUI

The microphone button sits in the composer's tool row, left of the model selector.

StateWhat you see
IdleGrey microphone
RecordingRed stop square, a live level meter, and a running m:ss clock
Silent inputThe meter stays flat and no sound — check your microphone appears
TranscribingRed microphone, transcribing…
FailureRed message beside the button (permission denied, no speech, engine error)

The level meter is the point, not decoration: a recording indicator that never moves cannot be told apart from a muted microphone. Fourteen bars track the incoming level, sampled from the same stream through an AnalyserNode, so you can watch your voice arrive before you stop. The analyser is deliberately never connected to the audio output — that would feed back through your speakers.

If the input stays silent for about 1.5 seconds, the meter dims and the hint appears. That covers the case the meter alone cannot: a muted microphone and a paused speaker both look flat, and the hint tells you which to check.

The transcript is appended to whatever the draft already contains, separated by a space. It is never sent for you.

Cleanup pass

Speech contains filler, false starts, and the occasional mishearing. After recognition, the transcript is passed through the model this deployment already selects, which removes filler (um, uh, a hesitant "like"), collapses stutters, adds punctuation, and fixes obvious mishearings — while being explicitly forbidden from changing meaning, translating, or touching identifiers and technical terms.

Worked example, on a deliberately filler-heavy recording:

StageText
WhisperUmm, so, uhh, like, can you please, you know, refactor the parser and then, umm, run the tests.
CleanedCan you please refactor the parser and then run the tests?

A small cleaned up marker appears beside the button afterwards, so you know a pass ran.

The cleanup is best-effort and can never cost you a transcript. Any of these keeps the recognizer's own text instead, with the reason written to the host log:

  • the cleanup call errors, times out, or returns nothing;
  • the rewrite changes length by less than half or more than 1.8× — the signature of a model answering the request rather than cleaning it;
  • no model route can be resolved.

Turn it off with polish: off, or pin it to a specific (small, cheap) route with polishProvider + polishModel so cleanup does not ride the large model the agent is using. The voice_transcribe tool accepts polish: false per call.

On any surface, as a tool

The same engine is also a model-facing tool, so a CLI, SDK, or ACP session can transcribe an audio file with no browser involved:

dsh --profile headless "transcribe /path/to/meeting.m4a and save the transcript next to it"

The tool is voice_transcribe:

ArgumentRequiredMeaning
pathyesAudio file to transcribe; relative paths resolve against the session working directory
languagenoLanguage code, or auto; defaults to the configured language
polishnofalse returns the recognizer's text without the cleanup pass

It returns the transcript as text, or a message naming the failure. Accepted containers are whatever ffmpeg decodes — wav, mp3, m4a, webm/opus, ogg, flac. The browser route and the tool share one engine, one cache, and one configuration.

Because the tool exists on every surface but the route only where a browser can reach it, the plugin declares no hard dependency on webServer: a headless or SDK profile loads it, registers the tool, and skips the route.

Configuration

Every option has a working default, so configuration is optional. To change one, override the row in $DSH_HOME/profiles/web/cordis.patch.yml — a patch replaces a row's entire config, so restate every key you still want:

- id: voice-input
  name: dsh-voice-input
  config:
    model: small.en
    language: en
    computeType: int8
    timeoutMs: 300000
KeyDefaultMeaning
modelbase.enWhisper model size or name
languageenSpoken language, or auto to detect
computeTypeint8CTranslate2 compute type
timeoutMs300000Deadline for one transcription
polishconservativeClean the transcript with a model; off returns the raw text
polishTimeoutMs15000Deadline for the cleanup call alone
polishProvider / polishModelcurrent selectionPin cleanup to a specific route
pythonPathpackage .venvInterpreter with faster-whisper
scriptPathbundled CLIThe transcription script
cacheDir$DSH_HOME/cache/voice-modelsModel weights cache

The cleanup uses the model the deployment already selects, so it needs no second credential and no extra configuration. It is a normal ctx.llm.stream call — the same route the agent uses — and it runs with temperature: 0 under its own deadline.

Choosing a model

ModelSizeNotes
tiny.en~75 MBFastest, least accurate
base.en~150 MBDefault; roughly 1 s of CPU per 5 s of speech
small.en~500 MBNoticeably better, about 3× the compute

Drop the .en suffix for multilingual models (base, small) and pair them with language: auto. A model that is not cached downloads on first use and needs network access once.

If huggingface.co is blocked, set DSH_VOICE_HF_ENDPOINT to a mirror such as https://hf-mirror.com before starting DSH.

Privacy

  • Audio is recorded by the browser and posted to the harness's own loopback HTTP server.
  • Transcription runs in a local child process. No audio, text, or model request leaves the machine; the only network access is the one-time model download.
  • The route is guarded by the composition's connection trust fence when one is mounted (Host/Origin checks that defeat DNS rebinding and cross-site posts), and by the server's loopback binding otherwise.
  • Recordings are staged under the OS temp directory and are not cleaned up automatically; clear %TEMP%\dsh-voice-input (Windows) or /tmp/dsh-voice-input when you want the space back.

How it works

Two halves, one npm package:

  • Host half (index.mjs) registers a model-facing voice_transcribe tool on tools, and one exact POST route, /voice-input/transcribe, on the composition's webServer. It stages audio to a temp file and runs the bundled CLI (lib/engine.mjs) through the shell service.
  • Browser half (client.cjs) registers a microphone button in the conversation.input.left slot, records with MediaRecorder, meters the same stream through an AnalyserNode, posts the bytes, and calls inputActions.setDraft() with the result.

Five details are load-bearing, and every one was found by testing rather than reading:

  • The host stages; the engine only reads. A confined shell refuses writes outside the session workspace — including the platform temp root — so an engine that wrote its own decoded copy would fail with PermissionError. Keeping every write in the host process keeps the engine usable under any sandbox policy.
  • The shell parses a command string with its own interpreter (pwsh -Command on Windows). Quoted paths in command position are therefore not an executable to PowerShell, so the Windows command is wrapped in cmd /c.
  • The route is deferred, not declared. webServer must not be a hard dependency (no headless/SDK/ACP profile has one), and it is published after this row's apply — so reading it once there sees undefined and the route silently never registers. A second row that waits for it is equally wrong: the boot audit fails on any entry left pending (N entry did not activate). ctx.inject(['webServer'], …) is the form that defers correctly and stays green everywhere.
  • The tool schema is standard JSON Schema. The provider validates parameters verbatim, so requiredness must be the object-level required: [...] array. The in-repo defineTool helper accepts a per-property required: true convenience form that is rejected on the wire.
  • A slot name is not a service. The browser Loader resolves a bundle's inject list before apply runs, and the client boot audit fails the page on any pending entry — so declaring conversation.input.left produced pending (waiting for service: conversation.input.left) and took the whole Web boot down. slots.inject(…) is what waits for the slot declaration; the bundle declares only slots. The timer service is probed rather than declared for the same class of reason, and the meter degrades harmlessly without it.
  • The resident worker reads a live pipe. The subprocess seam exposes a live child.stdout only for the stdio mode 'pipe'; a {maxBytes} spec routes the bytes to collected and leaves child.stdout undefined. The worker needs each protocol line the moment it flushes, so it spawns with 'pipe' — and because nobody else then reads a live stderr, it drains that too (an unread OS pipe fills at ~64 KiB and silently blocks the child forever). Pending requests are released on exit before any await, and run() enforces timeoutMs as a hard deadline, so a wedged request ends in a clean error naming the deadline instead of a hang that only killing DSH can clear.

The engine venv lives at the package root (<package>/.venv), resolved from lib/engine.mjs, so both host responsibilities find the same interpreter.

Troubleshooting

SymptomCause and fix
No microphone buttonThe bundle is installed in a different profile than the one booted. dsh web always means the web profile.
pending (waiting for service: webServer) at bootAn older build declared webServer as a hard dependency or as a separate waiting row. Update; the route is now deferred with ctx.inject.
Button appears but every recording fails with an engine path errorAn older build resolved the venv relative to the wrong module. Update, then re-run python/setup.py.
Invalid schema for function 'voice_transcribe'An older build used the in-repo per-property required: true form. Update.
faster-whisper is not installed in this interpreterRun python/setup.py; or point pythonPath at the interpreter you did install into.
ModuleNotFoundError: av / metadata_errorsav 19 removed an argument faster-whisper passes. setup.py pins av<19; reinstall with it.
Microphone permission was deniedAllow the microphone for the harness origin in your browser's site settings.
no speech was recognizedThe recording was silent or too short.
the transcription engine exited with code 1 … PermissionErrorAn older build staged the payload for the child to write. Update.
A transcription never returns; only killing DSH clears itAn older build spawned the worker with collected output, which the seam hides behind collected, so the pump read undefined and no request was ever answered or failed. Update, then restart DSH.

Development

Install the checkout directly into a scratch profile built from the web template:

dsh --profile voicetest --from-default-profile web --dump-config
dsh plugin --profile voicetest add /path/to/dsh-voice-input
dsh --profile voicetest

The engine can be exercised without the harness:

python/.venv/Scripts/python.exe python/transcribe.py --audio sample.webm

Contract tests (npm test) cover the host and browser plugin shapes, the row split, the bundle's registration format, and the tool's parameter schema. They cannot cover browser behaviour, so tools/cdp-check.mjs drives a real Chrome over the DevTools Protocol: it loads the running page, asserts the client entry applied, clicks the microphone, and samples the DOM to confirm the clock advances and the meter moves.

# Against a harness already listening (its logged URL carries the token):
node tools/cdp-check.mjs "http://127.0.0.1:3080/?token=<token>" 9333

It launches its own Chrome on a separate debugging port, uses Chrome's fake capture device, and exits non-zero if the boot audit fails or the meter or clock stays still. That is the half no unit test can reach — it is what caught the timer and inject defects.

License

MIT

Comments

Loading…

Similar plugins

dsh-voice-input

by QDchuan

Voice input for DeepSeek Harness: a microphone seat in the composer and a Web settings page that transcribes through the browser or any OpenAI-compatible /audio/transcriptions endpoint, with a fully o

Manifest valid

★ 0

MIT

JavaScript

Sep 12, 2026

dsh plugin --profile web add dsh-voice-input

by CharlesLiuZC

Voice-to-text plugin for DeepSeek Harness (desktop & web): composer mic button with cloud (SiliconFlow) or local offline (FunASR / faster-whisper) transcription.

Manifest valid

★ 0

TypeScript

Oct 2, 2026

dsh plugin --profile web add dsh-voice-context

by PerryLink

Voice-first session loop for DeepSeek Harness: a composer microphone button with browser/local speech-to-text (Web Speech, FunASR, whisper.cpp), a speak tool for text-to-speech replies (browser, edge-

Vision & MultimodalTools & CapabilitiesSecurity & AuditManifest valid

★ 16

↓ 1k/wk

Apache-2.0

TypeScript

Sep 25, 2026

dsh plugin --profile web add dsh-talk

by wangzhanchao883

Hold-to-talk voice input for the DeepSeek Harness web composer: hold the mouse on the input box, speak, release to insert the text into the draft. Local SenseVoice ASR via sherpa-onnx: no API key, off

Development & InfrastructureVision & MultimodalManifest valid

★ 3

↓ 363/wk

MIT

JavaScript

Oct 2, 2026

dsh plugin --profile web add dsh-hold-to-talk

by baisama-cloud

Speech-to-text voice input plugin for DeepSeek Harness (DSH) web GUI: click the mic in the composer to turn speech into text in the input box. Browser Web Speech API + OpenAI-compatible Whisper (OpenA

Vision & MultimodalTerminal & ClientsTools & CapabilitiesDevelopment & InfrastructureManifest valid

★ 3

MIT

JavaScript

Aug 31, 2026

dsh plugin --profile web add dsh-stt-input

by zhuiyueya

Voice for DeepSeek Harness(dsh) — speech-to-text input + read-aloud TTS for text-only DeepSeek, zero API key.

Manifest valid

★ 4

↓ 436/wk

MIT

JavaScript

Aug 15, 2026

dsh plugin --profile web add dsh-voice