@harness-flow/dsh-voice
dsh-voice — turn-based voice loop for DeepSeek Harness: pluggable Qwen / MiMo / local ASR+TTS engines, agent-driven speak/listen tools and browser PTT UI, built for interviewer presets
120 results
dsh-voice — turn-based voice loop for DeepSeek Harness: pluggable Qwen / MiMo / local ASR+TTS engines, agent-driven speak/listen tools and browser PTT UI, built for interviewer presets
Voice input for DeepSeek Harness: hold Space in any editable field to speak, release to insert the transcript at the caret. Zero dependencies, browser Web Speech API. / 语音输入:在输入框长按空格键说话,松开自动上屏。零依赖,基于浏览器 Web Speech API。
Reconciliation: statements, invoices and ledgers that have to balance. Bank/credit statement PDFs and invoice batches into ledger-ready tables, then reconciled — on a shared key in integer cents, or with no shared key at all by amount, date window, reference numbers and fuzzy counterparty names.
Editable local and desktop dictation for DeepSeek Harness
Workspace-scoped durable memory for DeepSeek Harness and optional voice integrations
DeepSeek Harness plugin: turn lines of narration into a finished vertical reel. Free neural voice-over, burned-in captions, no API keys.
QQ Bot (QQ 机器人) IM channel plugin for DeepSeek Harness (dsh): per-peer sessions, group @-mention and keyword triggers, image generation, TTS voice replies, attachment handling.
DSH KITT — Spanish-first voice for the DeepSeek Harness web UI: hands-free conversation, push-to-talk (Whisper), read-aloud replies, and a floating overlay window with global hotkeys. · DSH KITT 语音:为 DeepSeek Harness 打造的语音插件——免提对话、按键说话(Whisper)、朗读回复,以及带全局快捷键的悬浮窗;界面支持中文、英文和西班牙语。
Wake-word voice input for DeepSeek Harness: an always-on, low-power keyword spotter wakes the agent, then local ASR turns what you say next into a submitted message. No API key, no cloud, no typing.
Voice input plugin (dual-face): mic button beside the composer send action, live recording bubble streaming PCM to a local Qwen3-ASR python service managed by the host, plus an ASR environment settings page (clone runtime/model repos, create venv, run server). Bilingual UI (zh/en)
DeepSeek Harness plugin: when a run finishes or needs approval, get a link on your phone, hear the result from an ElevenLabs voice agent and say what to do next.
Codex realtime live voice for DeepSeek Harness: WebRTC + Frameless Bidi + session delegation.
Realtime Avatar (realtimeavatar.ai) for the DeepSeek Harness: the developer's API key held by the harness, the public docs as on-demand skills, rta_* tools over the public REST API, and a /rta command that walks a developer from no key to a live avatar call.
Voice for DeepSeek Harness backed by Xiaomi MiMo: browser-native 🎤/🔊 UI (MiMo TTS read-aloud) + voice_transcribe/voice_speak calling MiMo ASR/TTS directly, with a configurable voice map and in-conversation speech strips. Fork of zhuiyueya/dsh-voice (MIT).
文档文字识别。用户提供 PDF/扫描件/图片(合同、发票、书页、截图),需要提取文字、转成可编辑文本时使用。扫描件自动 OCR(macOS Vision,中英文)。Document OCR: extract editable text from PDFs, scans, and images (contracts, invoices, book pages, screenshots) via macOS Vision.
DSH plugin: compress verbose voice-dictation text into token-efficient prompts, fully local.
Provider-neutral full-duplex voice Agent capability for DeepSeek Harness.
开口即成文 · Speak-to-prompt for DeepSeek Harness:云端 ASR 语音识别 + 提示词优化 + 填入草稿/自动发送,跨平台 macOS / Windows。
Voice control for DSH web: speech-to-text into the composer (with auto-send) and spoken playback of assistant replies via the Web Speech API
Minimal Chinese voice input plugin for the DeepSeek Harness Web UI
Voice-feedback plugin for DeepSeek Harness: a speak tool, readReplies narration, an LLM verbalizer pass, per-session voices, chimes, a monitoring panel, and a native macOS floating pet. zh/en i18n.
A speaker button beside the Like button: read any assistant reply aloud with the browser's built-in speech engine, with speed and voice control. · 在点赞旁加一个小喇叭,朗读 DSH 的回复,支持倍速与音色。
DSH voice reader: read LLM replies aloud as they stream. Edge-tts built-in, any OpenAI-compatible cloud TTS plugs in via the settings panel, plus a voiced waiting-phrase ticker while the model thinks.
Voice input for DeepSeek Harness: dictation chunked by pauses and voice messages, each with its own provider fallback chain (Deepgram, Groq, HuggingFace, local whisper.cpp, plus any OpenAI-compatible API of your own).