Bundle
dsh-plugin-vision-toolkit
Vision toolkit for DeepSeek Harness -- glance, ground, detect, crop CLI tools for text-only agents to understand images
- Source
- YYTbit
- stars
- 1 stars
- License
- MIT
- Updated
- Updated 6 days ago
Readme
# dsh-plugin-vision-toolkit Vision toolkit for DeepSeek Harness -- give text-only agents the ability to see images. ## What it does Provides CLI tools that call a vision API (DeepSeek VL, GPT-4V, or any OpenAI-compatible endpoint) to describe, locate, detect, and crop elements from images. Registered as a dsh skill so agents know when and how to use them. ## Tools - `glance` -- describe, ask about, or OCR an image - `ground` -- locate a specific element (returns bounding box) - `detect` -- find all instances of an element kind - `crop` -- cut a region from an image ## Install ```sh dsh plugin --profile your-profile add dsh-plugin-vision-toolkit ``` ## Configuration Set environment variables: ```bash export VISION_API_KEY=sk-xxx # Vision API key (falls back to DEEPSEEK_API_KEY) export VISION_BASE_URL=https://... # API endpoint (falls back to DEEPSEEK_BASE_URL) export VISION_MODEL=deepseek-vl2 # Vision model name ``` ## Usage examples ```bash # Describe an image glance screenshot.png # Ask a question glance screenshot.png -q "What error is shown?" # OCR glance screenshot.png --ocr # Find a button ground screenshot.png "the login button" # Output: 450,820,620,870 # Find all buttons detect screenshot.png "buttons" # Crop a region crop screenshot.png 450,820,620,870 button.png ``` ## How it works The plugin registers a skill in the system prompt that teaches the agent about the vision tools. When the agent encounters an image (user pastes one, references a screenshot, etc.), it calls the appropriate CLI tool which: 1. Reads the image file 2. Encodes it as base64 3. Sends it to the vision API with a prompt 4. Returns the text response The agent never sees raw pixels -- it gets text descriptions it can reason about. ## Supported vision providers - DeepSeek VL (deepseek-vl2, deepseek-vl2.5) - OpenAI GPT-4V / GPT-4o - Any OpenAI-compatible multimodal endpoint ## License MIT -- YYTbit
Install
dsh plugin --profile web add github:YYTbit/dsh-plugin-vision-toolkit
Profile: web
With the hub plugin installed, ask your agent to install it by name — it resolves the same plan shown here.
dsh plugin --profile web add github:stvlynn/dsh.fish#path:packages/dsh-plugin-hub
install dsh-plugin-vision-toolkit from the hub
- This package builds from source on install. pnpm will ask you to allow its build script — that is permission to run the package’s code on your machine, outside the agent sandbox. Only allow sources you trust.
- This source has no pinned commit, so a later push upstream changes what installs. Prefer pinning a commit.