dsh-vision-local
Local-first vision for text-only DeepSeek Harness agents: route image understanding to a local OpenAI-compatible vision model, return structured JSON evidence (summary / OCR / layout / semantics / visual / uncertainty).
233 results
Local-first vision for text-only DeepSeek Harness agents: route image understanding to a local OpenAI-compatible vision model, return structured JSON evidence (summary / OCR / layout / semantics / visual / uncertainty).
Use a ChatGPT Codex subscription in DeepSeek Harness through standalone OAuth
Adaptive image routing for DeepSeek Harness: pass images directly to native multimodal models and transcribe them only for text-only models.
Multi-provider image generation and editing tools for DeepSeek Harness
dsh-eye-vision: give text-only DeepSeek Harness models eyes — image understanding via any OpenAI-compatible multimodal API. Fork of dsh-free-vision with custom-provider fixes and an allowed-directories whitelist.
DSH 对话内联生图插件:OpenAI 兼容接口,图片以 generated-image 块直接显示在对话框内,设置项在 设置→插件→可配置。
DeepSeek Harness plugin: a third-party image-generation relay. Adds the 生图配置 card (Settings → Plugins → 插件配置) for configuring an OpenAI-compatible (/v1/images/generations + /v1/images/edits), DashScope (通义万相 async-task), Replicate or fal.ai image API; registers image_generate (文生图/图生图/局部重绘/指令编辑), imagegen_status, imagegen_test, imagegen_configure tools; installs a 生图 skill so the model auto-triggers on 生图/画图 mentions.
Multimodal plugin for DeepSeek Harness: understand images and generate images through configurable OpenAI-compatible or DashScope endpoints.
Text-to-speech for DeepSeek Harness: speak agent replies in the Web UI with a provider fallback chain (OpenAI, ElevenLabs, Google, Azure, Groq, Deepgram, OpenRouter, Edge, Piper, eSpeak).
OpenAI Responses API-backed web search provider for dsh: registers into ctx.web (web_search tool with server-side retrieval), plus a Web settings card for endpoint/model/key. Hot-pluggable — mounted via profile bundle patch, no dsh source changes.
OpenAI Codex CLI integration for DeepSeek Harness: run one-shot codex tasks, repo reviews and session resumes as native dsh tools, with read-only sandbox by default.
Nous Portal OpenAI-compatible LLM adapter bundle for DeepSeek Harness
Multi-protocol web_search provider for DeepSeek Harness
Generate images and videos in DSH chat via OpenAI-compatible providers configured in Model settings; dedicated Settings menu and workspace media_gen output.
语音朗读(MiniMax / OpenAI 兼容 TTS,服务商与音色自选)——每条回复旁的朗读按钮、输入框自动朗读开关、设置页自选服务商与音色。零构建,纯 JS。
给 DSH agent 装一双眼睛:调用在线 VLM(多供应商、OpenAI 兼容)分析本地图片、URL 与会话上传的附件;勾选的模型特色能力会实时注入系统提示
Configurable OpenAI Responses-compatible native web search provider for DeepSeek Harness, with a Settings UI
DeepSeek Harness plugin: bridge image-bearing messages to a user-configured OpenAI-compatible vision endpoint when the active text model cannot see images
Image generation for DeepSeek Harness: a generate_image tool with pluggable providers — the FAL queue API or any OpenAI-compatible images API. The picture is shown inline in the conversation; the model receives either a link (works with any chat model) or the image itself (needs dsh-vision-bridge or a vision-capable model).
OpenAI Codex OAuth provider and native settings UI for DeepSeek Harness
DeepSeek Harness plugin: content-routing LLM adapter. Text-only requests go to DeepSeek; requests containing images go to a user-configurable OpenAI-compatible vision provider. Read_image works natively for text-only models.
A vision bridge for text-only DeepSeek Harness models: registers an `analyze_image` tool that answers questions about images via ANY OpenAI-compatible vision/multimodal endpoint (SiliconFlow, DashScope, Zhipu, OpenRouter, Ollama, ...).
Tolerate OpenAI-compatible model streams that end without a terminal finish_reason so DSH turns complete instead of aborting.
Token Rhythm (OpenAI Responses web_search) search provider for the DeepSeek Harness web capability seam (ctx.web)