All plugins

DSH / BUNDLE / CLIENT-UI

dsh-voice-live

v0.1.0tangzheng202202 / dsh-voice-live118f563533

InstallableBundlesUI & client pluginsCommunity · Topic auto-analysisWeb UI

Overview

dsh-voice-live

Real-time duplex voice for DeepSeek Harness: Volcengine streaming ASR/TTS, agent reply narration, barge-in, wake word, live captions. | 实时双工语音插件:火山流式 ASR/TTS、回复朗读、打断、唤醒词、实时字幕。

README / EN

Package documentation

Registry summary

Real-time duplex voice for DeepSeek Harness: Volcengine streaming ASR/TTS, agent reply narration, barge-in, wake word, live captions. | 实时双工语音插件:火山流式 ASR/TTS、回复朗读、打断、唤醒词、实时字幕。

dsh.pub verifies the pinned bundle contract, runtime facts, and distribution semantics. The complete README remains in the source repository.

Read the full README on GitHub

LIMITATIONS

Known limitations

- **Wake word** ships as a Web Speech API implementation (Chrome/Safari, zh-CN, keyword match on the recognized transcript). It is off by default (keeps the mic held). Two constraints are documented rather than papered over: - Chrome's Web Speech recognition routes through Google's cloud recognizer, which is unreachable from mainland networks — the settings probe (`检测唤醒可用性`) surfaces the exact `onerror` reason (e.g. `network`) instead of failing silently; Safari (Apple's service) is the working browser on macOS. - The preferred local offline engine, sherpa-onnx WASM keyword spotting, is blocked upstream: the `sherpa-onnx-wasm` npm package is deleted (404 on npm/jsdelivr/unpkg/npmmirror) and the current GitHub release assets are task-specific builds whose wasm binaries contain no KWS kernel (verified: no `sherpa_onnx_*_kws` symbols). vosk-browser was evaluated as a fallback but its fixed Chinese model vocabulary cannot constrain custom keywords such as the default wake phrase. The `WakeWordDetector` interface keeps the seam for a future local WASM engine once a distribution exists. - **Echo cancellation** relies on the browser's AEC (Chrome AEC3) applied to the WebRTC mic track. `VoiceController.getMicAec()` reports whether the browser actually applied AEC. `new Audio()` playback is not covered by Chrome's AEC (Chromium bug 687574), so the loop stays half-duplex: mic frames are gated while TTS plays, and barge-in stops playback before new recognition. The controlled echo test (speaker vs headphone, AEC on/off) is documented in the deliverable notes; the runtime decision is half-duplex + interrupt. - **Live captions** surface as the incremental composer draft (asr_partial); there is no floating caption overlay yet. - Server-side endpointing depends on Volcengine's optimized bidirectional endpoint (`bigmodel_async` + `enable_nonstream`); a plain `bigmodel` stream only returns partials and never finalizes on its own. - The package's `*.host.spec.ts` tests are excluded from the repository host aggregate's typecheck (TS6307 tripped by the aggregate excluding `packages/client/*/src/**`); they run under vitest and the host half builds through the package `tsc -b`.