dsh-multimodal
Multimodal eyes and hands for DeepSeek Harness: vision transcription, OCR, and text-to-image via OpenAI-compatible backends, with an in-conversation generated-image card.
132 results
Multimodal eyes and hands for DeepSeek Harness: vision transcription, OCR, and text-to-image via OpenAI-compatible backends, with an in-conversation generated-image card.
Workspace-bound arbitrary file upload, reading, OCR, and rendering for DeepSeek Harness.
Matter-aware legal workspace dashboard and document agent tools for DeepSeek Harness
Host-level vision bridge for text-only models: analyze_image tool (Ollama local / Xiaomi MiMo cloud / any OpenAI-compatible endpoint) returning structured evidence.
dsh plugin: recognize attached images locally with Tesseract OCR and send only the recognized text to the model. Text models never receive image bytes; vision passthrough is opt-in.
图片识别插件:自动判断当前模型是否具备视觉能力,有则用当前模型并按插件预设提示词分析,无则调用插件配置的视觉模型。文本模型可直接在对话框粘贴/上传/拖拽图片,发送时自动写入 DSH 附件存储(永久),消息区渲染缩略图,模型自动调用 image_vision / ocr / ground / crop 系列工具识别与精读;模型选择器保持简洁。
DeepSeek Harness 全能插件:识别/生图/改图一体化,无需切换模型——用常规 DeepSeek 模型即可自动调用视觉与生图模型(gemini_vision / gemini_generate_image / gemini_optimize_image)。多后端:Gemini 原生 + 任意 OpenAI 兼容服务(GPT-4o、Qwen-VL、GLM-4V、Moonshot、gpt-image、DALL-E、Flux、Stable Diffusion、OpenRouter、硅基流动、各类中转等),生成后自动视觉自检反馈,优于 modlens。
基于 DeepSeek-OCR1 光学压缩记忆系统:把记忆渲染为图像存储,支持 SoM 分段、年龄衰减/模糊化、激活召回、DSH Agent 检索
Eyes for text-only DeepSeek on DeepSeek Harness: image understanding via Antigravity IDE quota (default, flash/pro), OpenAI-compatible VLM endpoints, Gemini API, or local Ollama — model-invokable vision tool, wrapper adapters for deepseek/opencode-go, evidence memory, and a polished client panel.
DSH teacher plugin: Socratic tutor that leads you to answers from a markdown question set, tracks knowledge gaps in-session, and retests them on a spaced-repetition schedule.
DeepSeek Harness plugin that turns UI screenshots into implementation-grade web specs using OCR, deterministic geometry, scene graphs, assets, and render comparison.
Windows Computer Use for DeepSeek Harness: window-bound screenshots, robust OCR, verified clicks, pure-OCR mode, pluggable vision models.
DSH PaddleOCR (百度 PaddleOCR-VL 文档布局解析) plugin: OCR tools plus a settings card and task panel
DeepSeek Harness native plugin: vision capabilities for text-only LLMs (describe, OCR, VQA, layout analysis, clipboard)
DeepSeek Harness bundle for a Windows desktop UI context picker and MCP server for Codex, DeepSeek Harness, and AI agents
KiroCrew bridge for DeepSeek Harness: let your dsh agent delegate to a persistent, self-evolving KiroCrew workspace over ACP (JSON-RPC 2.0 over stdio).
Persistent vision plugin for DeepSeek Harness: registers the vision_analyze tool backed by a configurable multimodal model, with a settings-page UI. Zero dependencies.
Give DeepSeek Harness agents the ability to read images directly: a model-facing read_image tool that answers questions about an image through any OpenAI-compatible vision endpoint.
DeepSee: DeepSeek Harness vision, multi-model routing, and Gemini visual self-checks before delivery
DSH Plugins 4U 总插件:在设置中展示本仓库的自定义插件目录
macOS Vision OCR tool plugin for DeepSeek Harness: read text from screen capture, clipboard images, or image files (zh-Hans + en-US). Unofficial community project.
Native interactive visual-reasoning plugin for DeepSeek Harness: precise pixel grounding (SOM grid, zoom, annotate, measure, diff, color, OCR) + MiMo V2.5 multimodal backend, with zero external MCP servers.
WeChat Official Account content studio for DeepSeek Harness: anti-homogenization writing methodology, blessing-image visual baseline, gpt-image cover pipeline with OCR acceptance, interaction rules, and the measured xiaolvshu (newspic) draft web API.
Local vision 'eyes' for DeepSeek Harness (DSH): screen tool (capture screen or image -> local OpenAI-compatible VLM description) and ocr tool (Windows built-in OCR, zero model / GPU / cloud).