Bundle
@flowingboy/dsh-local-perf
Durable DeepSeek Harness bundle: the local-model performance tuning carried as a re-installable plugin layer (compaction, tool-result pruning, time context, cloud title routing, text-toolcall guard) — survives dsh updates
- Source
- flowingboy
- License
- MIT
- Updated
- Updated 6 days ago
Readme
# dsh-local-perf Durable DeepSeek Harness plugin bundle carrying the **local-model performance tuning process** as a re-installable layer — so it **survives dsh updates** instead of living in hand-edited patch files that a version bump can wipe or whose rationale dies with the author. Install once, re-apply forever: ```sh # from the GitHub repository git clone https://github.com/flowingboy/dsh-local-perf.git /tmp/dsh-local-perf dsh plugin --profile web add file:/tmp/dsh-local-perf # or from a checkout: dsh plugin --profile web add file:./dsh-local-perf ``` The bundle joins the profile's layer stack after `dsh-base` / `dsh-web-app`, re-applies all tuning rows on every boot, and vendors its own copy of the `text-toolcall-guard` plugin (self-contained, no dsh checkout required). --- ## The complete performance process (why every knob exists) Everything below was learned the hard way on an M5 Max running four local OpenAI-compatible servers (Ollama / MLX / Rapid-MLX / MLX-DSpark) against the DSH web GUI. The bundle encodes the *conclusions*; this README preserves the *reasoning*. ### Incident log | Date | Symptom | Root cause | Fix | | --- | --- | --- | --- | | 2026-08-18 | "fan spin / no response" | A 114,650-token prefill on Qwen3.5-122B — session context had grown unbounded because the web bundle disables auto-compaction | Re-enable `compaction-basic` at `thresholdRatio: 0.6` | | 2026-08-19 | session-title starves the interactive step | The title LLM request fires in the same second as a turn's first step; on the single-slot mlx-dspark server it queues behind the interactive generation and one of them starves past the idle budget | Route `session-title-llm` to the cloud model (`deepseek-official` / `deepseek-v4-flash`) | | 2026-08-19 | mlxdspark timeouts with tools present | The mlx-dspark server buffers the whole generation and emits no data events until it finishes; its 15s SSE keepalive comments are discarded by the OpenAI SDK parser and never reset the idle watchdog | Raise `timeoutMs` / `streamIdleTimeoutMs` to 600000 | | 2026-08-20 | tool calls appear as literal text | Local Qwen3.8-27B-8bit fell out of the structured `tool_calls` protocol under long tool-heavy steps and wrote `<tool_call>` prose the harness never executes | Ship the `text-toolcall-guard` plugin (vendored here) | | recurring | hallucinated "today" | No clock context in the prompt | Enable `time-context` (Asia/Shanghai, 10 min refresh) | ### Layer 1 — model configs (`settings.example.yaml` → `~/.dsh/settings.yaml`) Machine-specific (paths, ports, model ids), so the bundle carries them as a **template**, not a runtime patch. Copy the `llm-pi-ai` section into `~/.dsh/settings.yaml` on a fresh machine. The recurring principles: - **Timeout ≠ prefill tolerance.** A slow local server needs `timeoutMs` + `streamIdleTimeoutMs` ≥ worst-case prefill + reasoning + decode. Gemma 4 31B prefills at ~180 tok/s (system prompt + tool schemas ≈ 13k tokens → ~70s), mlx-dspark buffers whole generations, so both budgets sit at 300–600s. - **Retry only TRANSPORT.** Connection-level failures happen before prefill and are cheap; a TIMEOUT must never re-prefill a long prompt. - **Context window ≤ practical prefill budget.** 262144 tokens at ~180 tok/s is minutes of prefill. Lower to 32K–64K; compaction at 0.6× keeps sessions safely under the server limit. - **maxTokens ≤ decode budget.** At ~27 tok/s decode, 16K output is ~10 min. Cap at 8192–16384 so one step's worst case fits the timeout budgets. - **Reasoning effort default "off"** (or the server's lowest level) for quick, low-latency local loops; the local model's thinking stream still renders as a DSH reasoning block when enabled. ### Layer 2 — cordis rows (`cordis.patch.yml`) | Row | What | Why | | --- | --- | --- | | `time-context` | per-step clock | kills hallucinated dates | | `compaction-basic` | auto-compact at 60% | bounds prefill; the 08-18 incident fix | | `tool-result-pruner` | drop stale tool results | keeps them off later requests | | `command-compact` | manual `/compact` | escape hatch | | `session-title-llm` | title via cloud model | keeps the local slot free for the interactive step | ### Layer 3 — text-toolcall guard (`plugins/text-toolcall-guard`) Vendored from `@deepseek-ai/dsh-text-toolcall-guard` (built `lib/` + `src/`). When a step closes with no native tool calls but the assistant text carries `<tool_call>` / `<function=…>` markers at line starts, the guard steers a corrective message so the model re-issues the call natively. Bounded to `maxCorrections` (2) per turn per agent; marker detection requires line-start placement so prose that merely quotes the format is not corrected. Peers are resolved at runtime from the dsh installation's `profiles/node_modules` fallback (the designed out-of-tree-plugin path), so no registry fetch needed. --- ## Installation ```sh # from this directory dsh plugin --profile web add file:$(pwd) # verify the layer joined the stack dsh --profile web --dump-config | grep -A3 "dsh-local-perf" ``` The web profile's `cordis.patch.yml` should then only hold rows this bundle does NOT own (currently: none — everything moved into the bundle). ## Updating the bundle ```sh git pull # or edit locally cd plugins/text-toolcall-guard # rebuild the vendored guard if its src changed pnpm exec tsc -p tsconfig.json --outDir lib --declarationDir lib/types # reinstall the layer dsh plugin --profile web add file:$(pwd) ``` ## Publishing (GitHub / dsh-plugin ecosystem) Published at **https://github.com/flowingboy/dsh-local-perf** (public, `main`, topics: `dsh-plugin`). To re-publish after local edits: ```sh git add -A && git commit -m "dsh-local-perf: ..." git push origin main # topics (once) gh repo edit dsh-local-perf --add-topic dsh-plugin ``` ## Layout ``` cordis.patch.yml the perf layer (all tuning rows) settings.example.yaml model-config template (copy to ~/.dsh/settings.yaml) plugins/text-toolcall-guard/ vendored guard plugin (lib + src) README.md this document — the preserved process ```
Install
dsh plugin --profile web add github:flowingboy/dsh-local-perf
Profile: web
With the hub plugin installed, ask your agent to install it by name — it resolves the same plan shown here.
dsh plugin --profile web add github:stvlynn/dsh.fish#path:packages/dsh-plugin-hub
install flowingboy-dsh-local-perf from the hub
- This source has no pinned commit, so a later push upstream changes what installs. Prefer pinning a commit.