Bundle
@tkliuxing/dsh-hypatia
Long-term memory for DeepSeek Harness backed by Hypatia: a host-side CLI adapter, a plugin-owned SQLite control ledger, same-request recall, and narrow memory tools. Requires the `hypatia` CLI on PATH.
- Source
- tkliuxing
- License
- MIT
- Updated
- Updated 2 days ago
Readme
# dsh-hypatia
[中文文档](./README.zh.md)
Long-term memory for [DeepSeek Harness](https://github.com/deepseek-ai/deepseek-harness), backed by the [Hypatia](https://github.com/MarchLiu/hypatia) knowledge graph.
The plugin runs Hypatia itself, in host code. The model is not responsible for logging, database orchestration, permissions, retries, or deletion — it only proposes what is worth remembering.
What you get:
- **Recall in the same request** — relevant project memories are retrieved and attached to the turn that needs them, inside a fixed time and size budget, and always failing open
- **Exact project scoping** — memories belong to one project, derived from the canonical workspace path; cross-project leakage is prevented by a host-side ledger, not by hoping content tags line up
- **Verified writes** — every write is read back and compared before it counts as stored, so "saved" means saved
- **Two-stage forget** — you see exactly what will be deleted before anything is deleted, and cleanup status is reported honestly rather than optimistically
- **No Bash required** — memory works in `read-only` and `workspace-write` sessions, because the plugin never asks the model to shell out
## Prerequisites
**The `hypatia` command must be on your PATH,** and **Node 22.5+** (the control ledger uses `node:sqlite`). At load the plugin resolves the binary to an absolute path and checks its version; if either fails it logs a warning and memory stays inactive.
The adapter is written against the **hypatia 0.1.4** CLI contract and refuses anything older, because its output classification depends on observed per-command behaviour rather than exit codes. Override with `adapter.minVersion`, or set `adapter.requireVersionCheck: false` to proceed unverified.
```sh
git clone https://github.com/MarchLiu/hypatia
cd hypatia && cargo build --release
# put target/release/hypatia on your PATH
# Optional: the BGE-M3 embedding model, only needed for vector search
mkdir -p ~/.hypatia/default
hf download BAAI/bge-m3 --local-dir /tmp/bge-m3
cp /tmp/bge-m3/onnx/model.onnx ~/.hypatia/default/embedding_model.onnx
cp /tmp/bge-m3/onnx/model.onnx_data ~/.hypatia/default/model.onnx_data
cp /tmp/bge-m3/onnx/tokenizer.json ~/.hypatia/default/tokenizer.json
```
## Installation
```sh
# From npm
dsh plugin --profile web add @tkliuxing/dsh-hypatia
# Straight from GitHub (plain JS, no build step)
dsh plugin --profile web add github:tkliuxing/dsh-hypatia
# From a local path, for development against a checkout
dsh plugin --profile web add /path/to/dsh-hypatia
# When running dsh from a source checkout, use pnpm dsh instead:
pnpm dsh plugin --profile web add /path/to/dsh-hypatia
```
The published package is **`@tkliuxing/dsh-hypatia`**. The unscoped `dsh-hypatia`
name on npm belongs to this project's pre-rewrite release and is not updated.
Upgrading from an install made under that old name? Remove it first, or the
profile carries two entries for one plugin — and, because both resolve to the
same code, the plugin can load twice against one ledger:
```sh
dsh plugin --profile web remove dsh-hypatia
dsh plugin --profile web add @tkliuxing/dsh-hypatia
```
**Restart dsh** after installing or after editing `index.js`, `src/`, or `skills/`.
## Usage
Recall and summary ingestion are automatic. Beyond that, the agent has six tools it uses on your behalf:
| You say | What happens |
|---|---|
| "remember: this project forbids eval" | `memory_remember` stores one user-confirmed rule in this project's scope |
| "what do we know about the retry policy?" | `memory_search` returns this project's memories, labelled as reference data |
| "forget what you know about the old API" | `memory_forget_preview` shows the exact entries first; `memory_forget_confirm` deletes only what you approved |
| "did that actually save?" | `memory_status` reports verified, pending, and uncertain counts, plus how much of the project automatic recall scores |
| "settle whatever is still unverified" | `memory_reconcile` re-checks unverified operations against the knowledge base by stable key |
For knowledge-graph administration — shelves, archives, embedding models, export, or a deliberately unscoped search across the whole graph — the `hypatia` skill drives the CLI directly. That path does require `danger-full-access`.
## How it works
```text
DSH durable session log
|
| turn notifications, compaction summaries
v
dsh-hypatia host plugin
- memory authorization (independent of the file sandbox)
- project/scope derivation, provenance, stable operation IDs
- node:sqlite control ledger and retry queue
- recall cache, deadline, and context budget
|
| execFile(absoluteHypatiaPath, fixedArgv) shell: false
v
Unmodified Hypatia CLI
```
| Module | Responsibility |
|---|---|
| `src/policy.js` | memory capabilities, frozen at load |
| `src/identity.js` | project scope, stable names, operation IDs, provenance |
| `src/ledger/` | the plugin-owned SQLite control plane |
| `src/adapter/` | the one place a subprocess is spawned |
| `src/mutations.js` | intent → CLI → read-back verification → receipt |
| `src/recall.js` | same-request recall inside `agent/pre-step` |
| `src/retry-driver.js` | drains the retry queue inside the session that scheduled it |
| `src/tools.js` | the narrow `memory_*` tools |
| `src/ingest/` | idempotent ingestion of DSH compaction summaries |
[GOAL.md](./GOAL.md) is the authoritative architecture document, including the phases that are deliberately not implemented yet.
## Configuration
Everything is optional; override on the cordis row:
```yaml
- insert:
- id: dsh-hypatia
name: '@tkliuxing/dsh-hypatia'
config:
memory:
preset: standard # disabled | read-only-recall | standard | full
projectId: null # pin one scope across worktrees
state:
dir: ~/.dsh/dsh-hypatia
adapter:
shelf: default
timeoutMs: 10000
maxConcurrentReads: 1 # see "One process at a time" below
recall:
enabled: true
deadlineMs: 200
maxResults: 5
maxBytes: 10240
candidatePool: 50 # ledger records scored per turn
searchScanLimit: 200 # ledger records memory_search scans
hypatiaSupplement: true
vectorSupplement: false
ingest:
compaction: true
reconcile:
batchSize: 50 # operations and cleanups settled per pass
retryDriver: true # drain the retry queue in-session
```
### Coverage caps
Recall and `memory_search` both score a **newest-first** slice of the ledger, so a
project with more memories than the cap reaches the model only through the Hypatia
full-text supplement. Neither cap is silent: recall reports its ceiling once per
scope in the log, `memory_search` names it in the tool's `note`, and `memory_status`
returns `recall_coverage`. Raise `recall.candidatePool` to widen the pool — it costs
one wider SQLite read per turn and no subprocess.
### Memory authorization
Memory capabilities are **independent of the DSH file sandbox**. `read-only`, `workspace-write`, and `danger-full-access` govern what the *agent* may touch; they are not memory consent. Presets:
| Preset | Grants |
|---|---|
| `disabled` | nothing |
| `read-only-recall` | recall only |
| `standard` (default) | recall, semantic write, delete, reconcile |
| `full` | adds global-rule write and shelf administration |
Global-rule writes and transcript mirroring are never available to automatic paths, whatever the preset says.
## Limits worth knowing
These are deliberate, and the plugin reports them rather than papering over them.
- **One process at a time.** Every `hypatia` invocation opens all registered shelves and DuckDB takes an exclusive file lock, so concurrent invocations fail with `Conflicting lock is held` — measured at 3 failures out of 4 concurrent `hypatia query` calls against hypatia 0.1.4. The adapter therefore serializes every call, reads included. Raise `maxConcurrentReads` only if nothing else can touch the same shelves.
- **Deletion is scoped honestly.** Forget tombstones a record immediately, deletes it from the active shelf, and verifies absence. It cannot reach Hypatia exports, backups, other shelves, unknown user-created relations, or the DSH transcript — and reports `cleanup-uncertain` instead of claiming success when verification is incomplete.
- **A failing write retries three times, then stops.** Backoff is 1s, 5s, 30s; after that the operation is dead-lettered and `memory_status` counts it. Retries are drained in-session by an armed timer, not a poll, so a transient lock conflict settles without waiting for the next dsh start — but nothing retries forever, and a payload conflict is never retried at all.
- **Vector recall is off by default.** Hypatia's top-K cannot pre-filter by scope, so results must be over-fetched and filtered afterwards. Enable `recall.vectorSupplement` only after benchmarking your dataset.
- **Background extraction is not implemented.** GOAL.md gates it NO-GO until the Phase 0–2 fault and security tests pass; setting `extraction.enabled` logs a warning and changes nothing.
- **Full-transcript mirroring is not implemented.** It stays off until its consent, retention, and cleanup prerequisites exist.
## Performance
`npm run bench` measures the CLI against the configured recall deadline on a throwaway shelf it creates and removes. Re-measured on hypatia 0.1.4, Node 22.22.3, darwin/arm64:
| Records | Concurrency | full recall P50 | P95 | max | Within 200 ms deadline |
|---|---|---|---|---|---|
| 100 | 1 | 47.8 ms | 52.2 ms | 58.2 ms | yes |
| 100 | 4 | 97.8 ms | 189.2 ms | 190.0 ms | yes |
| 1,000 | 1 | 50.3 ms | 54.9 ms | 61.8 ms | yes |
| 1,000 | 4 | 100.4 ms | 198.9 ms | 199.1 ms | yes |
Serialized concurrency is the cost driver, not dataset size: ten times the records costs about 3 ms, while four concurrent sessions roughly quadruple the P95.
Read the last row carefully. It clears the deadline by about 1 ms, and the individual `jse query` and `fts search` operations behind it already miss it (P95 203.4 ms, max 207.6 ms). Recall fails open, so exceeding the deadline costs coverage rather than turns — but **four concurrent sessions at 1,000 records is the measured ceiling**, not a comfortable margin. Re-measure before raising `adapter.maxConcurrentReads` or planning for a larger corpus.
## Development
```sh
npm test # full suite
node --test tests/ledger.spec.js # one file
npm run bench -- --sizes 100,1000 # performance gate
```
`skills/` is self-maintained here — it was once synced from the hypatia repo, but the two are now decoupled. Edit `skills/*/SKILL.md` directly.
## The TRIGGER bridge is gone
Earlier versions injected `[hypatia-memory] TRIGGER:*` messages and asked the model to run `hypatia` through Bash. That mode has been **removed**: it wrote protocol text into the durable transcript, had no durable operation IDs or write receipts, could lose the final assistant reply, and overloaded `danger-full-access` as memory consent.
A profile that still sets `legacyBridge.enabled: true` loads normally and logs a warning naming the removal — the key does nothing and can be deleted. Everything it used to do is now the `memory_*` tools plus automatic recall, neither of which needs Bash or a full-access session.
## License
[MIT](./LICENSE)
Install
dsh plugin --profile web add github:tkliuxing/dsh-hypatia
Profile: web
With the hub plugin installed, ask your agent to install it by name — it resolves the same plan shown here.
dsh plugin --profile web add github:stvlynn/dsh.fish#path:packages/dsh-plugin-hub
install tkliuxing-dsh-hypatia from the hub
- This source has no pinned commit, so a later push upstream changes what installs. Prefer pinning a commit.