dsh-video-lens
Give text-only DeepSeek Harness agents video understanding: scene-aware frame sampling + VLM + optional ASR transcript fused into timeline evidence. / 给纯文本模型的视频理解插件(场景感知抽帧 + VLM + 可选语音转录)
35 results
Give text-only DeepSeek Harness agents video understanding: scene-aware frame sampling + VLM + optional ASR transcript fused into timeline evidence. / 给纯文本模型的视频理解插件(场景感知抽帧 + VLM + 可选语音转录)
DSH 专属语音输入插件:点按或按住 Alt 说话、松开/再点按转文字,支持热词替换表(hot.txt)、自定义润色提示词、录音电平指示。默认本地离线识别(SenseVoice,零配置零 key、音频不出本机),自动回退浏览器 Web Speech,可选云端 ASR 与润色(复用 DSH 模型)。Voice input for DeepSeek Harness: tap or hold Alt to talk, get text in the composer — local SenseVoice by default, zero config, zero API key.
Voice input plugin for DeepSeek Harness
Voice AI girlfriend for DeepSeek Harness: FunASR mic input, Qwen3-TTS voice replies, companion animation window, QQ two-way chat. Needs the repo's voice bridge + NapCat.
Voice input plugin for DeepSeek Harness: microphone → local/browser speech recognition → text submitted as a normal chat message (input-only, preset-agnostic)
Xiaomi MiMo search + multimodal tools for DSH agents: mimo_search/vision/audio/video/asr/tts.
Voice-first session loop for DeepSeek Harness: a composer microphone button with browser/local speech-to-text (Web Speech, FunASR, whisper.cpp), a speak tool for text-to-speech replies (browser, edge-tts, piper), event announcements with mute, and speak-to-interrupt.
Full-duplex voice mode for DeepSeek Harness: streamed ASR -> LLM -> TTS with barge-in
VocoType voice input bridge for DSH Web: mic button + recording panel in the composer, auto-insert recognized text (dedupe, auto-launch/deploy, mtime-optimized polling)
DSH web plugin: local offline FunASR voice input (paraformer int8 onnx sidecar, Web Speech fallback, LLM polish).
DSH WebUI 语音输入插件:输入框麦克风按钮(Alt+V)→ 录音 → 浏览器 Web Speech API 或本地 FunASR 后端转写 → 文本回填输入框(不自动发送)
DeepSeek Harness Web 语音输入插件:按住说话的麦克风按钮 + 本机离线 sherpa-onnx SenseVoice ASR
Context-aware voice input for DeepSeek Harness with Web Speech, local SenseVoice transcription, model polish, editable Composer drafts, and user-controlled sending
Voice input plugin for DeepSeek Harness web (China-ready): Alibaba Cloud DashScope ASR via a local bridge. Mic button in the composer, streaming recognition, cursor-aware insertion, silence auto-stop.
Local-first, full-duplex voice for DeepSeek Harness, orchestrated by Muxiva
Speech plugin for DeepSeek Harness: per-message speak button, composer voice input, and auto-announce toggle, over cloud TTS/ASR with Web Speech API fallback
DSH Web 本地离线语音输入插件:浏览器采集麦克风 → host 拉起本地 FunASR (SenseVoiceSmall) 识别 → 文字填入输入框。
Real-time duplex voice for DeepSeek Harness: Volcengine streaming ASR/TTS, agent reply narration, barge-in, wake word, live captions. | 实时双工语音插件:火山流式 ASR/TTS、回复朗读、打断、唤醒词、实时字幕。
聊天框语音输入按钮 for DeepSeek Harness: 点击麦克风说话,多引擎转写(智谱 GLM-ASR-2512 / 本地 faster-whisper / Gemini / OpenAI)自动填入输入框。一个按钮,所见即所得。
SenseVoiceSmall-powered local speech-to-text for the DSH Desktop input box: click the mic beside the composer, speak, and the recognized draft (with language/emotion tags) is written into the input.
Speech capability plugin for the DeepSeek Harness (dsh) web host: a token-gated /s/api route family serving audio transcription (ASR) and synthesis (TTS) over configurable providers
低成本视频理解工具:B站链接/BV/本地视频 → 信息层(ASR+场景+对象轨迹+YOLO)→ 摘要+问答。问题驱动动态路由分层(L0/L1/L2)、语义层复用、预算上限。引擎自包含,无需外部依赖。
SenseVoice 语音输入插件 for DeepSeek Harness:在对话输入框旁添加麦克风按钮,录音后调用本地 SenseVoice 服务转成文本填入输入框。首次使用自动下载模型并显示进度,后端由插件自动启动。
常用的五种语音识别,中文普通话、英语、日语、韩语、粤语,自动识别语种。