Bundle
dsh-touchstone
DSH 试金石:给「自改造」补上评测这一环——用一套金标准用例,把改动前 vs 改动后跑一遍、打分、对比,好就留、不好就撤。
- Source
- Hwayn-pixel
- License
- MIT
- Updated
- Updated 6 days ago
Readme
# dsh-touchstone > **The missing evaluation half of DeepSeek Harness's self-evolution.** > You changed something — did it actually get better? Run the same **golden cases** against the current config and a **candidate**, score them, and get a before→after report. **Keep it if it improved; revert if it did not.** [DeepSeek Harness](https://github.com/deepseek-ai/deepseek-harness) plugin · MIT · works on desktop **and** web [English](README.md) | [简体中文](README.zh-CN.md) --- ## Contents - [What it is](#what-it-is) - [Why you want it](#why-you-want-it) - [Quick start (5 minutes)](#quick-start-5-minutes) - [A worked example](#a-worked-example) - [The four ideas](#the-four-ideas) - [How to write checks](#how-to-write-checks) - [What a run does](#what-a-run-does) - [Settings and where data lives](#settings-and-where-data-lives) - [What it does *not* do (limits)](#what-it-does-not-do-limits) - [FAQ](#faq) - [Development](#development) --- ## What it is A **touchstone** for your agent. Whenever you make a change — add a house rule, swap a prompt, tweak a preset — instead of going on a hunch, it runs that change against the **same set of golden cases**, side by side with "before", and hands you a scorecard: ``` case current candidate change lead-with-answer 0% 100% +100% ``` and closes with an **evidence-based** verdict: **✅ Better, keep it.** / **❌ Regressed, revert.** / **➖ No real difference.** The score only speaks for *these* cases — **the final call is still yours.** ## Why you want it DSH's whole bet is "everything is a plugin": prompts, tools, presets, even the agent loop are meant to be swapped at will. And the official account names the missing piece: > DSH has already split the agent into locatable, replaceable components — **but it still lacks a complete learning loop: propose a change, and use Eval to judge whether that change actually worked.** The ecosystem already has **undo / rollback** (recover from a bad change) but **no evaluation** (measure before you commit). dsh-touchstone fills exactly that gap. It is deliberately thin: it only uses the public **`ctx.llm.stream()`** seam, never patches the harness, and depends on no internal services — so it runs on both DSH 0.1.x (web) and 0.2 (desktop). ## Quick start (5 minutes) 1. **Install** (pick a profile): ```shell dsh plugin --profile desktop add dsh-touchstone # desktop app dsh plugin --profile web add dsh-touchstone # web UI ``` Restart DSH afterwards. 2. **Open it**: Settings (desktop: bottom-left **account menu → 设置**, or `Ctrl + ,`) → click **🪨 试金石** on the left. 3. **Add a case** — under "金标准用例 (golden cases)", click **+ 加一条用例**: - **name**: for you, e.g. "lead with the answer". - **prompt**: the sentence you would send to the agent, e.g. "Introduce artificial intelligence in one sentence." - **check**: click **+ 加一条检查项**, choose **必须包含 (must contain)**, and type the word you expect. 4. **Add a candidate** — under "候选方案", click **+ 加一个候选**: - **name**, e.g. "add a prefix rule". - **prompt**: the text you want to try, e.g. "Every answer must begin with exactly the four characters 「笔记:」." - **mode**: `append` (on top of current) or `replace`. 5. **Run**: click **跑一遍 (Run)** next to the candidate. 6. **Read the report** — scroll to **对账 (Report)** for per-case scores and the verdict. > 💡 Fields **save on blur**: after typing, click elsewhere (or press Tab) and it saves automatically. ## A worked example **You want** to add an "opening format" rule, but you don't know if the model will obey it. | Field | Value | |---|---| | case · name | `opening format` | | case · prompt | `Introduce artificial intelligence in one sentence.` | | case · check | `must contain` → `笔记:` | | candidate · name | `prefix marker` | | candidate · prompt | `Every answer must begin with exactly 「笔记:」, then continue normally.` | | candidate · mode | `append` | Click **Run**. A few seconds later: ``` prefix marker current 0% → candidate 100% case current candidate change opening format 0% 100% +100% ✅ Better, keep it. ``` **It adds evidence to the sentence you otherwise would not have**: not "I think it works", but "under the current config the instruction had no effect; under the candidate it did." ## The four ideas | Idea | What | Note | |---|---|---| | **Case** | one prompt + some checks | a question you want the agent to stay solid on | | **Check** | an assertion about the output | must contain / must not contain / regex / LLM judge | | **Candidate (variant)** | a prompt you want to try | append or replace; the control is the built-in "current (no change)" | | **Report** | candidate mean − baseline mean | per-case delta + overall verdict | ## How to write checks | Kind | Decision | When to use | |---|---|---| | **must contain** | the output contains this text | cheapest and most reproducible; "a word / conclusion / format must appear" | | **must not contain** | the output does **not** contain this text | prohibitions, e.g. "no 'Sorry'" | | **regex** | the output matches | format checks, e.g. `^笔记:` | | **LLM judge** | a second model call answers `PASS` / `FAIL` + one line | subjective criteria, e.g. "conclusion first, then reasons" | > **Rule of thumb:** prefer rule checks over the LLM judge — rule scores are the most trustworthy and reproducible; the judge is also a model, with its own bias. > Checks can carry a **weight** (default 1); the score is a weighted pass rate. ## What a run does 1. Take every **enabled** case × the candidates you ticked (plus the always-on control "current"). 2. For each pair, build one model call — `system` = the candidate's prompt, `messages` = the case prompt — and collect text and usage via `ctx.llm.stream()`. 3. Score: rule checks **in-process**; the LLM judge makes **one more model call**. 4. Store the run and return the report. The whole thing **only reads config and calls the model** — it never touches your files. ## Settings and where data lives - **Which model to run** — configure `provider` / `model` / `temperature` / `maxOutput` for `dsh-touchstone` in DSH settings, plus optional **judge** model (`judgeProvider` / `judgeModel`; blank = same as the runner). - **Data** — `$DSH_HOME/touchstone/bench.json` (`DSH_HOME` defaults to `~/.dsh`); stores cases, candidates, and the last **40 runs**. - To reset, delete that file (it is recreated). ## What it does *not* do (limits) - This is a **text-level** evaluation: it compares "the model's output under the same prompt", **not** a full agent loop. - Best for: **changing a prompt / a house rule / a preset's wording.** - Not for: "changing a tool's implementation" that needs real side effects — it gives no evidence there. - **The judge is also a model** and has bias; rule checks are the most reproducible. Treat the score as evidence, not a verdict. - The score only speaks for **your cases** — write them poorly and you measure the wrong thing. ## FAQ **Q: How long / how expensive is a run?** A: One model call per case × candidate (plus one more if you use an LLM judge). Few cases, short outputs → fast and cheap. Start with one or two cases. **Q: Can I compare two candidates instead of candidate-vs-current?** A: Each candidate is compared against "current". For candidate-vs-candidate, read the two reports side by side (same cases, same baseline). **Q: Why did my input box garble Chinese?** A: An early development build had this bug — the input posted on every keystroke and wrote the server value back, which fights the IME. **This version is fixed**: text edits stay local and save on blur. **Q: Will it touch my files?** A: No. It only reads config and calls the model. ## Development ```shell npm install npm run build # src/index.ts -> lib/index.js ; src/client/index.ts -> lib/client.js npm run typecheck npm test # pure-logic unit tests (run against the built bundles; no DSH needed) ``` - **Host half** (`src/index.ts`): storage + run engine + routes + settings schema. - **Browser half** (`src/client/index.ts`): the "试金石" settings page. - Local iteration: `node build.mjs --watch`, then refresh the page (client half reloads instantly). --- ## Authors & license - **Author: Hwayn (幻弈)** - **Collaborator: Yucheng Xiao (肖宇成)** — direction, requirements, testing - License: MIT
Install
dsh plugin --profile web add github:Hwayn-pixel/dsh-touchstone#f5248d6928d5399df2fb656c8cff530c5b50ee70
Profile: web
With the hub plugin installed, ask your agent to install it by name — it resolves the same plan shown here.
dsh plugin --profile web add github:stvlynn/dsh.fish#path:packages/dsh-plugin-hub
install dsh-touchstone from the hub
- This package builds from source on install. pnpm will ask you to allow its build script — that is permission to run the package’s code on your machine, outside the agent sandbox. Only allow sources you trust.