Skip to content
dsh.fish
Bundle

dsh-reset-handoff

Standard handoff protocol so DSH never restarts itself: a host plugin that files a JSON reset request for an external ops agent (e.g. Hermes) to execute preflight-snapshot → restart → health-check → recover, then delivers the result back to the requesting session.

Source
nicecx
License
MIT
Updated
Updated 6 days ago

Readme

# dsh-reset-handoff

> **DSH never restarts itself.** A host plugin that hands reset requests to an external ops agent over a versioned JSON protocol — preflight-snapshot → restart → health-check → recover — then delivers the result back to the requesting session after reboot.

## Why

A long-lived DeepSeek Harness (DSH) instance needs to restart for many reasons: reload plugins/config, apply settings, recover from a wedged state. But the **agent inside DSH should not restart DSH itself**:

- restarting kills the very process that issued it, so the agent has no chance to see the outcome;
- the agent cannot see what business is running (other live sessions, the relay channels, pending jobs);
- if the reboot fails, nobody is left to diagnose and recover.

The safe pattern is a **handoff**: DSH writes a *request*, a separate, independent ops agent (here: Hermes Agent) reads it, runs the restart with preflight/health/recovery, and writes a *result* that DSH reads back after it comes up.

## How it works

```
[DSH]  agent calls reset_handoff(reason)
         │  writes request.json (JSON protocol)
         │  (optional) triggers the external executor
         ▼
[ext]  ops agent reads request.json
         │  1. preflight — snapshot live sessions, relay state, pending jobs
         │  2. GATE     — pre-restart maturity gate (see below)
         │  3. restart  — restart the dsh web service (macOS launchd)
         │  4. health   — poll http://127.0.0.1:3080 until 200 (with timeout)
         │  5. recover  — verify relay/auth-proxy self-heal, list interrupted sessions
         │  6. result   — write result.json (status done/failed + per-stage detail)
         ▼
[DSH]  after reboot, the plugin reads result.json and delivers a readable
         summary back into the requesting session (followup), so the agent
         that asked can resume its interrupted work.
```

### Pre-restart maturity gate

The reference executor refuses to restart unless it is safe to do so. It checks:

1. **No pending approvals/questions** — relay `pending.json` has an empty `pending` list (a restart would otherwise drop the approval stack).
2. **Enough free disk** — at least `MIN_FREE_DISK_MB` (default 500 MB).
3. **Cooldown** — at least `RESTART_COOLDOWN_SEC` (default 60 s) since the previous restart, to break crash loops.

If any condition fails, the executor writes `result.json` with `status: "failed"`, `restart: { ok: false, gated: true }`, and a `gate` array listing each failed condition with its detail — **and does not restart**. The requesting agent (or user) sees the exact reason and decides when it is safe to retry.

## Tools

| Tool | Purpose |
| --- | --- |
| `reset_handoff(reason, scope?)` | Submit a reset request to the external ops agent. Never restarts DSH in-process. |
| `reset_status()` | Query the latest request and its result (read-only). |

Both tools are registered host-wide, so **every session's agent** can call them when a reset is needed.

## Protocol (v1)

The plugin and the executor are **decoupled** — they only share two JSON files under `~/.dsh/reset-handoff/` (override with `DSH_RESET_HANDOFF_DIR`):

**`request.json`** (written by DSH):

```json
{
  "schema": "dsh-reset-handoff/request",
  "version": 1,
  "id": "<uuid>",
  "requestedAt": "2026-08-30T12:00:00+08:00",
  "reason": "重新加载插件配置",
  "sessionId": "<requesting session id>",
  "requester": "dsh-reset-handoff",
  "scope": { "restartDshWeb": true, "healthCheck": true, "recoverInterrupted": true }
}
```

**`result.json`** (written by the executor):

```json
{
  "schema": "dsh-reset-handoff/result",
  "version": 1,
  "requestId": "<uuid>",
  "status": "done",
  "startedAt": "...",
  "finishedAt": "...",
  "preflight": { "liveSessions": ["..."], "relay": { }, "hermesJobs": ["..."] },
  "restart": { "ok": true },
  "health": { "ok": true, "checks": [ { "name": "dsh-web http :3080", "ok": true, "detail": "200" } ] },
  "recovery": { "resumed": ["..."], "report": "..." },
  "recoveryAction": {
    "ok": true,
    "attempts": [ { "attempt": 1, "restart": { "ok": true }, "time": "..." } ],
    "diag": { "keyErrors": [], "logTail": "..." }
  },
  "gate": [ { "name": "relay 无待审批/待回答诉求", "ok": true } ]
}
```

**Executor recovery contract** (the part that makes "recover DSH itself" real): if the health check fails after restart, the executor **must** attempt recovery, not just report failure:

1. **Diagnose** — read the dsh web error log tail and extract key errors (loader failures, missing deps like undici, `EADDRINUSE` multi-instance, OOM).
2. **Retry** — restart up to `MAX_RESTART_ATTEMPTS` (default 3) times with a cooldown between attempts.
3. **Observe** — after each restart, wait an initialization window (default 120 s) before judging success.
4. **Report** — write the outcome in `recoveryAction` (`attempts` + `diag`), so the requesting agent and the human see *why* it failed and how many tries were made.

Any executor that reads/writes these two files can drive the reset — Hermes, a custom script, a cloud function. The protocol is the contract.

## Install

```sh
dsh plugin --profile <profile> add github:<owner>/dsh-reset-handoff
```

Optional executor trigger: configure `triggerCommand` so `reset_handoff` also wakes the external agent (default: none — the executor may poll `request.json` instead). See `cordis.patch.yml` for the config shape.

## Executor (Hermes example)

A reference executor is included under `hermes/reset_agent.py` (pure Python, no deps). It is meant to live inside a Hermes profile (`reset-agent`) and be triggered by `hermes cron run <job>`:

```sh
python3 reset_agent.py            # run the five-step flow
python3 reset_agent.py --dry-run  # print the flow, don't restart
```

## Requirements

- DeepSeek Harness with the web profile (host plugin).
- The external executor must be able to restart the dsh web service (macOS `launchctl kickstart -k com.dsh.web`, or equivalent for your OS/init).

## Ops guardrails (read this before restarting anything)

Learned the hard way from a real 7-hour restart loop (2026-08-30). These rules are mandatory for any agent that manages a DSH host:

1. **Never create suicide/unconditional restart jobs.** No `launchctl submit` jobs containing `kickstart -k`, no `kill -9` on the DSH port, no unconditional restart logic. Restart only via the `reset_handoff` tool (which goes through the executor's gate) or DSH's own mechanism.
2. **Verify plugin dependencies before restart.** The DSH loader resolves from `~/.dsh/profiles/web/node_modules` — a missing transitive dep (e.g. undici) makes the whole plugin tree fail to load. Confirm deps exist and the tree loads cleanly before restarting.
3. **Health-check first, observe after.** Before any restart: `curl` the port, check for single instance (`lsof -i :3080`). After restart: wait a 2-minute observation window and confirm the PID is stable before proceeding.
4. **Make plugins degrade gracefully.** Missing config / bad fields should fall back to defaults with a friendly error, so calling agents never feel the need to edit plugin source or kill services to work around bugs.

## License

MIT

Install

dsh plugin --profile web add github:nicecx/dsh-reset-handoff

Profile: web

  • This source has no pinned commit, so a later push upstream changes what installs. Prefer pinning a commit.
Source