@liustack/modlens
Plug-in vision for text-only LLMs, powered by the free Antigravity CLI
127 results
Plug-in vision for text-only LLMs, powered by the free Antigravity CLI
Eyes for text-only DeepSeek Harness agents: built-in free vision chain (no key) + pixel-level vision tools (Q&A, grounding, crop, pixel diff, colors, OCR, SVG trace, cutout, screenshots). One-command install, no Python, image turns work like ordinary tool-calling turns.
DeepSeek Harness plugin: dispatch work to DSH agents from Claude Code / Codex, as native subagents with live progress
Give text-only DeepSeek Harness agents video understanding: scene-aware frame sampling + VLM + optional ASR transcript fused into timeline evidence. / 给纯文本模型的视频理解插件(场景感知抽帧 + VLM + 可选语音转录)
Self-contained DeepSeek Harness plugin for Provider login, model switching, image fallback, usage analytics, and same-port Web restart
Near-native image understanding for text-only DeepSeek Harness models
让 DeepSeek Harness 里任何纯文本模型都能读图。识图是按需调用的 tool —— 图片不进主模型上下文,不看就不花钱;附 23 处缺陷的评测集,换模型可自测。
Unified image understanding plugin for DeepSeek Harness (DSH). Visual twin adapter for native thumbnails + auto-analysis on any text-only model (incl. pi-ai providers); privacy/smart/strict routing; local tools (scan/OCR×4 engines (windows/macos/paddle/rapid)/crop/palette/compare/batch); document-to-image (pdf/word/excel/ppt); local photo editing (image_edit: resize/rotate/filter/composite/watermark/background-remove/upscale etc, pure CPU); optional external VLM bridge.
DeepSeek Harness plugin bundle: qwen_vision (Qwen-VL image understanding) and qwen_generate (Qwen-Image text-to-image and image editing) tools for text-only models
Visual media plugin for DeepSeek Harness: copy native image descriptions and securely normalize, inspect, and play scene-aware videos.
Auditable vision and cross-platform Computer Use runtime for DeepSeek Harness with source-preserving evidence.
DeepSeek Harness plugin: configurable vision model with vision_read_image tool, composer-bar vision-model selector, and automatic image-to-text conversion for text-only main models.
MiniMax multimodal bridge for DeepSeek Harness (DSH). One mmx_bridge tool covers describe/image/video/speech/music/cover/search/quota; optional web_search/read_image takeover; built-in client enhancement renders inline players/previews plus a settings-page management card in the Web GUI.
Per-model capability declaration for DeepSeek Harness: reasoning-effort levels (with wire spellings) and request modalities (text/image) for OpenAI-compatible providers — one settings section, no YAML hand-editing.
Free vision plugin for DeepSeek Harness (dsh): image understanding for text-only models with free-tier providers (Qwen3-VL-Flash / DeepSeek-OCR / Doubao). 免费视觉插件:纯文本模型看图能力,优先免费模型(通义千问 / 硅基流动 / 豆包)。
Eyes for text-only DeepSeek on DeepSeek Harness: image understanding via Doubao Web by default (zero cost, no API key), Antigravity IDE quota (flash/pro), Gemini API, or Cockpit proxy — model-invokable vision tool, wrapper adapters, evidence memory, and a bilingual client panel.
On-demand vision for text-only DSH sessions: images become markers, and a vision_describe tool sends only image + question to an OpenAI-compatible vision model
Xiaomi MiMo search + multimodal tools for DSH agents: mimo_search/vision/audio/video/asr/tts.
DSH Web plugin that lets text-only models see images: paste images in the chat and send without file paths; the model discovers its own vision tools. Multimodal models pass through natively.
Automatic Web image-to-text bridge plus vision and OCR tools for DeepSeek Harness
Guide Dog for DSH, powered by MiniMax — multimodal plugin: image/video/music/speech generation, vision inspection tools, voice mode, microphone voice input and real-time voice call mode.
Multi-agent collaboration suite for DeepSeek Harness: a user-configured specialist roster with on-demand dispatch (team_call / roundtable), model comparison, and a multimodal vision bridge — models come from the official provider flow, no bundled adapters.
Global model request headers plus image input, reasoning, and DeepSeek system-role compatibility for custom providers
让纯文本主模型(DeepSeek V4 等)也能接收图片附件:抹除纯文本路由的模态声明放行 0.1.1 准入门禁,图片投影为携带完整 attachmentId 的占位文本,由主模型委托视觉子代理经 read_image 读取。