Bundle
dsh-reconnect
Safe model-request retry for DeepSeek Harness with exponential backoff, model-availability recovery, and configurable unknown-error limits.
- Source
- MistRain-1
- stars
- 3 stars
- License
- MIT
- Updated
- Updated 4 days ago
Readme
# dsh-reconnect
> Safe model-request retry for DeepSeek Harness with exponential backoff, relay/proxy failure recovery, and a unified settings panel.
**Version:** `2.0.0`
- [中文文档 (Chinese README)](./README_CN.md)
## Project Description
`dsh-reconnect` is a Host-side plugin for DeepSeek Harness (DSH). It retries failed model requests so temporary network failures do not interrupt an agent turn or a long-running task.
The plugin listens to the DSH `agent/request-error` waterfall event. After the built-in `normal` Provider retry policy gives up, or when a Provider's `always` policy delegates recovery downstream, the Agent loop sends the same request again according to this plugin's policy.
## Suitable Use Cases
- Unstable relay stations, reverse proxies, API gateways, or forwarding services
- Relay nodes that disconnect, reset connections, or return incomplete responses
- Intermittent network failures and transport errors
- Temporary provider outages, server-side 5xx errors, rate limits, and timeouts
- Long-running tasks that should survive short-lived provider or network failures
## Retry Policy
| Condition | Behavior |
| --- | --- |
| `EMPTY_RESPONSE` / `RATE_LIMIT` / `SERVER` / `STREAM_CLOSED` / `TIMEOUT` / `TRANSPORT` | Retry indefinitely (transient connection/service failures) |
| Model missing or not configured (`MODEL_NOT_FOUND`, `MODEL_NOT_CONFIGURED`, or matching Provider text) | Retry indefinitely while waiting for model/account configuration to recover |
| `QUOTA` (insufficient balance or exhausted quota) | Do not retry by default; retried when `retryQuota: true` |
| Unknown errors such as `PI_AI_ERROR` | Retry indefinitely by default; disable `retryUnknown` to apply the configurable consecutive-failure cap |
| Tool/argument/unknown-tool/auth/credential/context-overflow/aborted | Never retried; step ends immediately |
Key design points:
- Only retries failures that "resending the same request" can fix — dropped connections, rate limits, timeouts, 5xx.
- Tool execution errors (`tool/result`) are outside this plugin's retry boundary. Error routing uses the machine code from `agent/request-error`, rather than guessing from message text; an unknown machine code uses the configured unknown-error path (bounded only when `retryUnknown` is disabled).
- Permanent errors (credentials, request content, context overflow) stop immediately; missing or unconfigured models are the deliberate exception and keep retrying while configuration recovers.
- `retryQuota` and `retryUnknown` control this plugin's fallback chain. The Host `always` policy delegates recovery to downstream waterfall listeners, so this plugin handles the recoverable failure categories once and does not create a second parallel retry loop.
## Backoff and Cancellation
- Exponential backoff: `1s -> 2s -> 4s -> ...`, capped at `60s` by default (the settings panel offers 1/2/5/10/30/60/120 seconds or a custom value; `maxDelayMs` overrides in milliseconds)
- Honors a positive `providerRetryAfterMs` when available. It is a provider minimum wait and is not capped by the local exponential-backoff limit; only the Node timer maximum applies.
- Logs the provider, error code, turn, step, retry count, and delay
- Emits standard `llm/retry` events so the Harness conversation UI shows the continuous retry count and countdown
- Stops when the current turn is aborted, the plugin is stopped, or the plugin is hot-reloaded; active waits are drained during cleanup.
## Visual Configuration
Open Settings → Plugin configuration. ReConnect is shown as a plugin card in that page:
- **Max wait per attempt**: drop-down presets (1/2/5/10/30/60/120 seconds) or a custom value
- **Retry on quota**: toggle, off by default
- **Retry unknown errors indefinitely**: toggle, on by default
- **Max retries for unknown errors**: default 3. Missing or unconfigured model errors are not subject to this cap and always retry.
Changes take effect immediately without restarting DSH.
## YAML Configuration
The plugin accepts optional values in its Cordis row config:
- `maxDelayMs` (integer milliseconds, default `60000`): caps the local exponential-backoff delay; for example `15000` produces `1s -> 2s -> 4s -> 8s -> 15s -> 15s...` indefinitely. A positive Provider `Retry-After` remains authoritative.
- `retryQuota` (boolean, default `false`): whether `QUOTA` errors are also retried.
- `retryUnknown` (boolean, default `true`): whether unknown errors are retried indefinitely.
- `unknownMaxRetries` (integer, default `3`): bounded cap for consecutive unknown failures in the same model step; missing or unconfigured model errors do not use this cap.
The settings service persists the plugin card values. **Restore defaults** removes the user overrides and immediately returns to the schema/Cordis defaults; a restart is not required.
The plugin's own retry policy is the downstream recovery for both Provider `normal` and `always` policies. The two layers do not multiply retries: `normal` schedules only its configured codes before this plugin, while `always` delegates the decision downstream.
`maxDelayMs` caps the local exponential wait, **not** a Provider `Retry-After` value and not the total retry duration.
```yaml
- insert:
- id: reconnect
name: dsh-reconnect
config:
maxDelayMs: 15000
retryQuota: false
retryUnknown: true
```
Invalid or missing `maxDelayMs` values fall back to `60000`.
## Download
The package is platform-independent. Choose one of the following methods.
### Method 1: GitHub Web Download
1. Open https://github.com/MistRain-1/dsh-reconnect.
2. Select **Code** and then **Download ZIP**.
3. Extract the archive and use the extracted `dsh-reconnect-main` directory as the plugin package.
### Method 2: Git Clone
Windows PowerShell:
```powershell
git clone https://github.com/MistRain-1/dsh-reconnect.git "$HOME\dsh-reconnect"
```
macOS:
```bash
git clone https://github.com/MistRain-1/dsh-reconnect.git "$HOME/dsh-reconnect"
```
Linux:
```bash
git clone https://github.com/MistRain-1/dsh-reconnect.git "$HOME/dsh-reconnect"
```
### Method 3: Download the ZIP from a Terminal
Windows PowerShell:
```powershell
Invoke-WebRequest -Uri https://github.com/MistRain-1/dsh-reconnect/archive/refs/heads/main.zip -OutFile dsh-reconnect.zip
Expand-Archive -Path dsh-reconnect.zip -DestinationPath .
```
macOS:
```bash
curl -L https://github.com/MistRain-1/dsh-reconnect/archive/refs/heads/main.zip -o dsh-reconnect.zip
unzip dsh-reconnect.zip
```
Linux:
```bash
wget https://github.com/MistRain-1/dsh-reconnect/archive/refs/heads/main.zip -O dsh-reconnect.zip
unzip dsh-reconnect.zip
```
After downloading, register the package in the DSH Cordis composition:
```yaml
- insert:
- id: reconnect
name: dsh-reconnect
```
A DSH restart is required after installing the persistent plugin package.
## Requirements
- DSH Host plugin
- Uses the `agent/request-error` waterfall event and the Host settings service
- Client side provides a unified-format settings card and never reads credentials
## License
MIT
Install
dsh plugin --profile web add github:MistRain-1/dsh-reconnect
Profile: web
With the hub plugin installed, ask your agent to install it by name — it resolves the same plan shown here.
dsh plugin --profile web add github:stvlynn/dsh.fish#path:packages/dsh-plugin-hub
install dsh-reconnect from the hub
- This source has no pinned commit, so a later push upstream changes what installs. Prefer pinning a commit.