Skip to content
dsh.fish
Bundle

dsh-gpu

GPU-aware execution layer for DeepSeek Harness: gpu_status / gpu_exec / gpu_run_bg tools, per-step GPU context injection, automatic CUDA_VISIBLE_DEVICES card selection

Source
zytsyj
stars
3 stars
License
MIT
Updated
Updated 4 days ago

Readme

# dsh-gpu

GPU-aware execution layer for [DeepSeek Harness](https://github.com/deepseek-ai/deepseek-harness) (dsh). Out-of-tree plugin; no harness patches required.

Agents get three tools — `gpu_status`, `gpu_exec`, `gpu_run_bg` — plus an optional per-step GPU context line. Cards are selected automatically (freest first) with `CUDA_VISIBLE_DEVICES` set in the command environment; pin a card explicitly when you care.

```
8 GPU(s), free: [0,1,2,3,4,5,6,7]
GPU0 Tesla V100-SXM2-32GB: 4264/32768MiB 0%util 40C
...
[gpus 1 — GPU 1 (auto: freest 1)] exit 0
```

## How it works

- **`gpu_status`** — one query, every device: memory used/total, SM utilization, temperature, and a free/busy verdict. A device is *busy* at or above 80% memory used or 50% utilization (both configurable).
- **`gpu_exec`** — one-shot command with a selected card: `CUDA_VISIBLE_DEVICES=<freest>` is passed through the mounted `ctx.shell` executor's environment. Auto-select or pin `gpuIndex`; select `count` cards for multi-GPU commands.
- **`gpu_run_bg`** — long-running GPU jobs (training, inference servers, benchmarks) register as a `gpu` job in `ctx.jobs`: returns a job id immediately, read with `job_output`, stop with `job_kill`.
- **Per-step context** (optional, on by default) — injects a one-line GPU snapshot into eligible steps (the `time-context` pattern), rate-limited to one sample per minute.

All execution rides the **mounted shell executor**. Local host, or any remote execution world (e.g. an SSH provider plugin) — dsh-gpu doesn't know or care where the GPUs are; it queries and launches through the same seam the `bash` tool uses.

## Install

dsh-gpu is an out-of-tree bundle plugin. Install and activate it in a profile with the official plugin command:

```bash
dsh plugin --profile <name> add dsh-gpu
```

The package's bundled `cordis.patch.yml` registers the plugin automatically. To override its configuration, add an entry with the same `id` to the profile's `cordis.patch.yml`:

```yaml
- insert:
    - id: gpu
      name: dsh-gpu
      config:
        stepContext: true
```

Load order note: place it after your execution-world plugins (e.g. an SSH provider) so the shell seam it queries is the one you intend.

## Configuration

```yaml
- id: gpu
  name: dsh-gpu
  config:
    stepContext: true      # per-step GPU snapshot line (default true)
    refreshIntervalMs: 60000  # min spacing between injected snapshots
    queryTimeoutMs: 10000     # nvidia-smi timeout
    busyMemoryPct: 80         # >= this % memory used => busy
    busyUtilPct: 50           # >= this % SM util => busy
```

## Notes & gotchas

- `nvidia-smi` **ignores** `CUDA_VISIBLE_DEVICES` — it always reports physical indices. `gpu_exec` selection still works as intended for CUDA programs; just don't use nvidia-smi output inside `gpu_exec` to verify the pinning.
- Selection is advisory, not a reservation: two concurrent agents can still pick the same card. For exclusive claims, pin `gpuIndex` from a `gpu_status` read in the same step.
- `gpu_run_bg` requires the jobs service in the composition (`@deepseek-ai/dsh-jobs` + `@deepseek-ai/dsh-tool-jobs`), the same dependency background `bash` has.
- Hosts without NVIDIA GPUs: `gpu_status` reports a clean `no-gpu` result instead of failing.

## Development

```bash
pnpm install
pnpm typecheck   # tsc --noEmit
pnpm test        # vitest unit and plugin lifecycle tests
pnpm build       # tsdown -> lib/
pnpm check:package  # publint + Are the Types Wrong
node tests/live-v100.mjs   # optional live probe (edit SSH target first)
```

Test fixtures are recorded from a live 8× Tesla V100-SXM2-32GB host (including one occupied card) — no mocking of nvidia-smi output formats.

## License

MIT

Install

dsh plugin --profile web add github:zytsyj/dsh-gpu#bd74d556b3ccc527989117bd850eab18971c7014

Profile: web

  • This package builds from source on install. pnpm will ask you to allow its build script — that is permission to run the package’s code on your machine, outside the agent sandbox. Only allow sources you trust.
Source