Bundle
dsh-reset-handoff
Standard handoff protocol so DSH never restarts itself: a host plugin that files a JSON reset request for an external ops agent (e.g. Hermes) to execute preflight-snapshot → restart → health-check → recover, then delivers the result back to the requesting session.
- Source
- nicecx
- License
- MIT
- Updated
- Updated 6 days ago
Readme
# dsh-reset-handoff
> **DSH never restarts itself.** A host plugin that hands reset requests to an external ops agent over a versioned JSON protocol — preflight-snapshot → restart → health-check → recover — then delivers the result back to the requesting session after reboot.
## Why
A long-lived DeepSeek Harness (DSH) instance needs to restart for many reasons: reload plugins/config, apply settings, recover from a wedged state. But the **agent inside DSH should not restart DSH itself**:
- restarting kills the very process that issued it, so the agent has no chance to see the outcome;
- the agent cannot see what business is running (other live sessions, the relay channels, pending jobs);
- if the reboot fails, nobody is left to diagnose and recover.
The safe pattern is a **handoff**: DSH writes a *request*, a separate, independent ops agent (here: Hermes Agent) reads it, runs the restart with preflight/health/recovery, and writes a *result* that DSH reads back after it comes up.
## How it works
```
[DSH] agent calls reset_handoff(reason)
│ writes request.json (JSON protocol)
│ (optional) triggers the external executor
▼
[ext] ops agent reads request.json
│ 1. preflight — snapshot live sessions, relay state, pending jobs
│ 2. GATE — pre-restart maturity gate (see below)
│ 3. restart — restart the dsh web service (macOS launchd)
│ 4. health — poll http://127.0.0.1:3080 until 200 (with timeout)
│ 5. recover — verify relay/auth-proxy self-heal, list interrupted sessions
│ 6. result — write result.json (status done/failed + per-stage detail)
▼
[DSH] after reboot, the plugin reads result.json and delivers a readable
summary back into the requesting session (followup), so the agent
that asked can resume its interrupted work.
```
### Pre-restart maturity gate
The reference executor refuses to restart unless it is safe to do so. It checks:
1. **No pending approvals/questions** — relay `pending.json` has an empty `pending` list (a restart would otherwise drop the approval stack).
2. **Enough free disk** — at least `MIN_FREE_DISK_MB` (default 500 MB).
3. **Cooldown** — at least `RESTART_COOLDOWN_SEC` (default 60 s) since the previous restart, to break crash loops.
If any condition fails, the executor writes `result.json` with `status: "failed"`, `restart: { ok: false, gated: true }`, and a `gate` array listing each failed condition with its detail — **and does not restart**. The requesting agent (or user) sees the exact reason and decides when it is safe to retry.
## Tools
| Tool | Purpose |
| --- | --- |
| `reset_handoff(reason, scope?)` | Submit a reset request to the external ops agent. Never restarts DSH in-process. |
| `reset_status()` | Query the latest request and its result (read-only). |
Both tools are registered host-wide, so **every session's agent** can call them when a reset is needed.
## Protocol (v1)
The plugin and the executor are **decoupled** — they only share two JSON files under `~/.dsh/reset-handoff/` (override with `DSH_RESET_HANDOFF_DIR`):
**`request.json`** (written by DSH):
```json
{
"schema": "dsh-reset-handoff/request",
"version": 1,
"id": "<uuid>",
"requestedAt": "2026-08-30T12:00:00+08:00",
"reason": "重新加载插件配置",
"sessionId": "<requesting session id>",
"requester": "dsh-reset-handoff",
"scope": { "restartDshWeb": true, "healthCheck": true, "recoverInterrupted": true }
}
```
**`result.json`** (written by the executor):
```json
{
"schema": "dsh-reset-handoff/result",
"version": 1,
"requestId": "<uuid>",
"status": "done",
"startedAt": "...",
"finishedAt": "...",
"preflight": { "liveSessions": ["..."], "relay": { }, "hermesJobs": ["..."] },
"restart": { "ok": true },
"health": { "ok": true, "checks": [ { "name": "dsh-web http :3080", "ok": true, "detail": "200" } ] },
"recovery": { "resumed": ["..."], "report": "..." },
"recoveryAction": {
"ok": true,
"attempts": [ { "attempt": 1, "restart": { "ok": true }, "time": "..." } ],
"diag": { "keyErrors": [], "logTail": "..." }
},
"gate": [ { "name": "relay 无待审批/待回答诉求", "ok": true } ]
}
```
**Executor recovery contract** (the part that makes "recover DSH itself" real): if the health check fails after restart, the executor **must** attempt recovery, not just report failure:
1. **Diagnose** — read the dsh web error log tail and extract key errors (loader failures, missing deps like undici, `EADDRINUSE` multi-instance, OOM).
2. **Retry** — restart up to `MAX_RESTART_ATTEMPTS` (default 3) times with a cooldown between attempts.
3. **Observe** — after each restart, wait an initialization window (default 120 s) before judging success.
4. **Report** — write the outcome in `recoveryAction` (`attempts` + `diag`), so the requesting agent and the human see *why* it failed and how many tries were made.
Any executor that reads/writes these two files can drive the reset — Hermes, a custom script, a cloud function. The protocol is the contract.
## Install
```sh
dsh plugin --profile <profile> add github:<owner>/dsh-reset-handoff
```
Optional executor trigger: configure `triggerCommand` so `reset_handoff` also wakes the external agent (default: none — the executor may poll `request.json` instead). See `cordis.patch.yml` for the config shape.
## Executor (Hermes example)
A reference executor is included under `hermes/reset_agent.py` (pure Python, no deps). It is meant to live inside a Hermes profile (`reset-agent`) and be triggered by `hermes cron run <job>`:
```sh
python3 reset_agent.py # run the five-step flow
python3 reset_agent.py --dry-run # print the flow, don't restart
```
## Requirements
- DeepSeek Harness with the web profile (host plugin).
- The external executor must be able to restart the dsh web service (macOS `launchctl kickstart -k com.dsh.web`, or equivalent for your OS/init).
## Ops guardrails (read this before restarting anything)
Learned the hard way from a real 7-hour restart loop (2026-08-30). These rules are mandatory for any agent that manages a DSH host:
1. **Never create suicide/unconditional restart jobs.** No `launchctl submit` jobs containing `kickstart -k`, no `kill -9` on the DSH port, no unconditional restart logic. Restart only via the `reset_handoff` tool (which goes through the executor's gate) or DSH's own mechanism.
2. **Verify plugin dependencies before restart.** The DSH loader resolves from `~/.dsh/profiles/web/node_modules` — a missing transitive dep (e.g. undici) makes the whole plugin tree fail to load. Confirm deps exist and the tree loads cleanly before restarting.
3. **Health-check first, observe after.** Before any restart: `curl` the port, check for single instance (`lsof -i :3080`). After restart: wait a 2-minute observation window and confirm the PID is stable before proceeding.
4. **Make plugins degrade gracefully.** Missing config / bad fields should fall back to defaults with a friendly error, so calling agents never feel the need to edit plugin source or kill services to work around bugs.
## License
MIT
Install
dsh plugin --profile web add github:nicecx/dsh-reset-handoff
Profile: web
With the hub plugin installed, ask your agent to install it by name — it resolves the same plan shown here.
dsh plugin --profile web add github:stvlynn/dsh.fish#path:packages/dsh-plugin-hub
install dsh-reset-handoff from the hub
- This source has no pinned commit, so a later push upstream changes what installs. Prefer pinning a commit.