Bundle
dsh-qwen-multimodal
DSH bundle: Qwen multimodal bridge — vision (qwen3-vl), speech-to-text (qwen3-asr), text-to-image (qwen-image), via the deepseek-vision skill scripts
- Source
- wuwangmao
- License
- MIT
- Updated
- Updated 3 days ago
Readme
# dsh-qwen-multimodal
A DSH bundle that gives text-only main models (e.g. DeepSeek) **three multimodal skills in one plugin** through Qwen APIs: vision, speech-to-text, and text-to-image — with a built-in generate-then-verify quality loop.
| Tool | Capability | Backend |
|---|---|---|
| `describe_image` | Image / screenshot / OCR / chart understanding (multiple images at once) | Qwen VL (default `qwen3-vl-flash`) |
| `transcribe_audio` | Speech / recording transcription (wav/mp3/m4a/aac/flac/ogg/amr) | Qwen3-ASR (`qwen3-asr-flash`) |
| `generate_image` | Generate images from text and save them locally | Qwen-Image (`qwen-image-plus`) |
Media never enters the main model context: visual/audio content is converted to text, and generated images are saved to local files with paths returned by the tool.
## How it works
All API calls reuse the original Python scripts in `skills/deepseek-vision/scripts/*.py` and the `.env` configuration (vision/audio use the Alibaba Cloud Bailian OpenAI-compatible endpoint; image generation uses the native multimodal-generation endpoint). The plugin itself is a pure-JS Cordis bundle depending only on the host's mounted `subprocess` / `tools` services — **no build step required for git installs**.
## Install
### From GitHub
```sh
dsh plugin --profile demo add github:wuwangmao/dsh-qwen-multimodal
```
### Local checkout / tarball
```sh
dsh plugin --profile demo add ./dsh-qwen-multimodal
# or
pnpm pack # then
dsh plugin --profile demo add ./dsh-qwen-multimodal-0.1.0.tgz
```
Before first use, configure your API key: copy `skills/deepseek-vision/.env.example` to
`skills/deepseek-vision/.env` and fill in `VISION_API_KEY` (create one in the Alibaba Cloud Bailian console; new users get free quota, college students get a ¥300 annual voucher). Vision/audio/image reuse the same key by default, or configure them separately (see `.env.example`).
### Python
**Python 3.10+ is required** (the image-generation script uses `int | None` type-annotation syntax).
`python` is resolved from the system `PATH` by default. If it cannot be resolved, restate the plugin row in your profile's `cordis.patch.yml` and set `config.pythonPath`.
### Configuration overrides
The plugin uses the bundled skill directory by default. To point it at an external directory (e.g. to reuse an existing `.env` and scripts, or to keep your key outside `node_modules`), restate the row:
```yaml
- insert:
- id: qwen-multimodal
name: dsh-qwen-multimodal
config:
skillDir: 'D:/qwen-vision'
pythonPath: 'C:/path/to/python.exe'
```
## Usage
Once loaded, the model can call the three tools directly:
- `describe_image({ images: ['screenshot.png'] })` — verbatim extraction of text/code/errors in images
- `describe_image({ images: ['chart.png'], prompt: '逐字提取图中所有文字,保留原样' })` — custom prompt
- `transcribe_audio({ audios: ['recording.m4a'], language: 'zh' })` — specify language for accuracy
- `generate_image({ prompt: 'a cute orange cat on a windowsill watching the sunset', out_dir: './out' })` — generate and save locally
- `generate_image({ prompt: '...', out_dir: './out', verify: true })` — generate, then automatically re-check the result with Qwen VL against the prompt (quality loop)
## Layout
```
dsh-qwen-multimodal/
├── package.json # dsh.bundle manifest
├── cordis.patch.yml # bundle layer: inserts the plugin row
├── src/index.js # plugin: registers the three model tools (pure JS)
├── scripts/selfcheck.mjs # self-check: node scripts/selfcheck.mjs
└── skills/deepseek-vision/ # skill assets: SKILL.md + Python scripts + .env.example
```
## License
MIT
Install
dsh plugin --profile web add github:wuwangmao/dsh-qwen-multimodal#ada1b30ba4f1ddc5d4a959796e1d5efac0cb0bda
Profile: web
With the hub plugin installed, ask your agent to install it by name — it resolves the same plan shown here.
dsh plugin --profile web add github:stvlynn/dsh.fish#path:packages/dsh-plugin-hub
install dsh-qwen-multimodal from the hub