Bundle
dsh-image-reader
Give DeepSeek Harness agents the ability to read images directly: a model-facing read_image tool that answers questions about an image through any OpenAI-compatible vision endpoint.
- Source
- zcXie777
- stars
- 2 stars
- License
- MIT
- Updated
- Updated 10 days ago
Readme
# dsh-image-reader
Give a text-only [DeepSeek Harness](https://github.com/deepseek-ai/deepseek-harness) agent the ability to **read images directly**: one model-facing `read_image` tool that asks any OpenAI-compatible vision endpoint about an image by its workspace path.
## Why
DeepSeek Harness is "everything is a plugin". This bundle mounts a single tool so the model can look at a screenshot, diagram, or photograph and answer questions about it, instead of only ever reasoning over text.
## Verification status
- Verified locally: `npm run typecheck`, `npm run build`, and `npm test` (16 tests) all pass.
- Not yet verified: a real end-to-end read against a live vision endpoint. The request/response logic is covered by a mocked-fetch unit test, but the plugin has not been smoke-tested inside a running dsh profile against a real multimodal model. Do that once with a real `VISION_API_KEY` before relying on it.
## Install
```sh
git clone https://github.com/zcXie777/dsh-image-reader.git
cd dsh-image-reader
npm install
npm run build # lib/ is not committed; build once after cloning
cd ..
dsh plugin --profile web add "$PWD/dsh-image-reader"
dsh plugin --profile headless add "$PWD/dsh-image-reader"
dsh --profile web --dump-config | grep image-reader
```
Restart a running Web profile after installing.
## Configure
`provider.baseUrl` and `provider.model` are required; the plugin never assumes a vendor. Override them in the profile patch row with the same id:
```yaml
- id: image-reader
config:
provider:
baseUrl: https://api.openai.com/v1
model: gpt-4o-mini
apiKeyEnv: VISION_API_KEY
lang: zh
timeoutMs: 60000
maxImageBytes: 10485760
allowedDirs: []
```
Set the key in the environment before starting the profile:
```sh
export VISION_API_KEY=sk-...
```
## Use
In a conversation, point the model at an image path and ask:
```text
read_image image="screenshot.png" query="What error is shown in this dialog?"
read_image image="diagram.png"
```
## Configuration fields
| Field | Default | Contract |
|---|---|---|
| `provider.baseUrl` | — (required) | OpenAI-compatible `chat/completions` base URL |
| `provider.model` | — (required) | Multimodal model name |
| `provider.apiKeyEnv` | `VISION_API_KEY` | Environment variable holding the API key |
| `lang` | `zh` | Answer language: `zh` or `en` |
| `timeoutMs` | `60000` | Whole-request deadline, 1000–600000 ms |
| `maxImageBytes` | `10485760` | Encoded-byte limit per image |
| `allowedDirs` | `[]` | Extra realpath-resolved input roots; the workspace is always allowed |
## Security
- Inputs resolve against the workspace and `allowedDirs` through `realpath`, so a symlink cannot escape the fence.
- Images are size-limited and extension-checked before upload.
- The key is read from the environment per call, never stored in config.
## Development
```sh
npm install
npm run typecheck
npm run build
```
## Publish
Tag the repo with the [`dsh-plugin`](https://github.com/topics/dsh-plugin) topic so it is discoverable, and publish to npm when ready.
## License
MIT
Install
dsh plugin --profile web add github:zcXie777/dsh-image-reader
Profile: web
With the hub plugin installed, ask your agent to install it by name — it resolves the same plan shown here.
dsh plugin --profile web add github:stvlynn/dsh.fish#path:packages/dsh-plugin-hub
install dsh-image-reader from the hub
- This package builds from source on install. pnpm will ask you to allow its build script — that is permission to run the package’s code on your machine, outside the agent sandbox. Only allow sources you trust.
- This source has no pinned commit, so a later push upstream changes what installs. Prefer pinning a commit.