Skip to content
dsh.fish
Bundle

dsh-tool-eyes

Local vision 'eyes' for DeepSeek Harness (DSH): screen tool (capture screen or image -> local OpenAI-compatible VLM description) and ocr tool (Windows built-in OCR, zero model / GPU / cloud).

Source
go-farther-and-farther
stars
2 stars
License
MIT
Updated
Updated 10 days ago

Readme

# dsh-tool-eyes

Local vision "eyes" for [DeepSeek Harness](https://github.com/deepseek-ai/deepseek-harness) (DSH).

Give your text-only agent eyes with **two model-facing tools**:

- **`screen`** — capture the screen (or describe an existing image file) through a
  local OpenAI-compatible vision endpoint (llama.cpp with `--mmproj`, LM Studio,
  Ollama, ...) and return the vision model's text description.
- **`ocr`** — extract **ALL text verbatim** with the **Windows built-in OCR
  engine**: zero model, zero GPU, zero cloud, milliseconds.

```
screen  =  screenshot / image  ->  local VLM  ->  text description (understanding)
ocr     =  screenshot / image  ->  Windows OCR -> verbatim text          (extraction)
```

## Why

DeepSeek's chat-completions line is text-only. Instead of switching your whole
conversation to a vision model, keep the text brain and add eyes as tools:

- **Private by default** — point `screen` at a local endpoint and images never
  leave your machine.
- **Cheap** — a 0.8B–4B local VLM is plenty for describing screens; `ocr` costs
  nothing at all.
- **Honest by default** — the `screen` prompt tells the VLM to *describe only
  what is visible* and never guess app/game/character names unless confirmed by
  on-screen text (this measurably cuts small-model name hallucination).

## Requirements

- Windows 10/11 (the `ocr` tool uses WinRT OCR; `screen` uses .NET for capture)
- Node.js >= 22.19, DeepSeek Harness >= 0.1.0-rc.6
- `screen` additionally needs any OpenAI-compatible VLM endpoint, e.g.:
  - llama.cpp: `llama-server -m model.gguf --mmproj mmproj.gguf --port 1235`
  - LM Studio (loaded vision model), Ollama, or any OpenAI-compatible gateway

## Install

> This package is published on **GitHub only** (not on npm).

```sh
dsh plugin --profile web add https://github.com/go-farther-and-farther/dsh-tool-eyes
```

Then restart `dsh web`. The `screen` and `ocr` tools appear in the agent's
toolkit automatically.

### Manual install (offline / from source)

Copy this package into the profile's `node_modules`, then register it in
`$DSH_HOME/profiles/<profile>/cordis.patch.yml`:

```yaml
- insert:
    - id: tool-eyes
      name: 'dsh-tool-eyes'
      config:
        baseUrl: http://127.0.0.1:1235/v1
        model: ''
        timeoutMs: 180000
```

## Configuration

Plugin config (all optional):

| key | default | meaning |
|---|---|---|
| `baseUrl` | `http://127.0.0.1:1235/v1` | OpenAI-compatible endpoint for `screen` |
| `model` | `''` | model id to send; empty lets the server decide (llama.cpp serves one model) |
| `timeoutMs` | `180000` | hard cap for one capture call |
| `captureScript` | bundled `capture.ps1` | override path to an alternate capture script |

Override in your profile's `cordis.patch.yml` (id-targeted):

```yaml
- id: tool-eyes
  name: 'dsh-tool-eyes'
  config:
    baseUrl: http://127.0.0.1:1235/v1
    model: qwen3.5-4b
    timeoutMs: 120000
```

## Usage

In a conversation, the agent can now:

- `screen` — "what is on my screen?", "describe this image file", with an
  optional `prompt` to focus on a region or detail.
- `ocr` — "read all the text on screen", "transcribe this error dialog".

Both accept an optional `image` path; without it they capture the screen.
The bundled PowerShell scripts can also be run standalone:

```powershell
powershell -NoProfile -ExecutionPolicy Bypass -File lib\capture.ps1 -Prompt "..." -BaseUrl http://127.0.0.1:1235/v1
powershell -NoProfile -ExecutionPolicy Bypass -File lib\ocr.ps1 -Image C:\path\x.png
```

## Privacy

- `ocr` is fully local (WinRT OCR, no network).
- `screen` sends the captured image to the configured `baseUrl`. Point it at a
  local endpoint (llama.cpp / LM Studio / Ollama) to keep images on your machine.

## Related

- [dsh-vision-proxy](https://github.com/Flyvhidbwo/dsh-vision-proxy) — automatic
  transcription of **attached images** in the chat input (Chatbox-style), so you
  don't need to give file paths. Pairs well with this plugin.

## Development

```sh
npm test    # node --test tests/
```

## License

MIT

Install

dsh plugin --profile web add github:go-farther-and-farther/dsh-tool-eyes

Profile: web

  • This source has no pinned commit, so a later push upstream changes what installs. Prefer pinning a commit.
Source