Skip to content
dsh.fish
Bundle

dsh-continual-harness

Standalone DeepSeek Harness plugin: continual harness self-evolution. The agent persists and refines reusable prompt notes, memories, skill contracts, and subagent specs through small evidence-backed edits, with automatic refinement gates, rollback, and prompt injection.

Source
jasen215
stars
9 stars
License
MIT
Updated
Updated 10 hours ago

Readme

# dsh-continual-harness

English | [中文](docs/readme/README.zh.md)

<p>
  <a href="LICENSE"><img src="https://img.shields.io/badge/license-MIT-blue.svg" alt="MIT License"></a>
  <a href="https://www.npmjs.com/package/dsh-continual-harness"><img src="https://img.shields.io/npm/v/dsh-continual-harness?cacheSeconds=86400" alt="npm version"></a>
  <img src="https://img.shields.io/badge/node-22+-339933.svg" alt="Node Version">
  <img src="https://img.shields.io/badge/typescript-6.0+-3178C6.svg" alt="TypeScript">
  <a href="https://github.com/jasen215/dsh-continual-harness/actions/workflows/ci.yml"><img src="https://github.com/jasen215/dsh-continual-harness/actions/workflows/ci.yml/badge.svg" alt="CI"></a>
  <a href="https://www.npmjs.com/package/dsh-continual-harness"><img src="https://img.shields.io/npm/dm/dsh-continual-harness?cacheSeconds=86400" alt="npm downloads"></a>
</p>

A **DeepSeek Harness (DSH) plugin for self-improving AI agents**, providing continual learning through persistent memory, periodic review and refinement, cross-session knowledge sharing, and automatic rollback on failure. It forms a closed loop of plan → validate → apply → rollback.

The design is inspired by the open-source [prime-agent](https://github.com/PrimeIntellect-ai/prime-agent) from Prime Intellect, a self-improving coding harness.

## Capabilities

A single npm package (`dsh-continual-harness`) takes effect through the following extension points once mounted:

| Capability | Mechanism |
| --- | --- |
| State projection (inject harness context each step) | `agent/pre-step` waterfall listener; incremental injection when the content digest changes |
| Review and automatic refinement | `session/event` listener on turn interval / compaction end; runs LLM review → plan → apply automatically |
| Manual refinement tool | Registers the `harness_refine` tool (directly callable by the LLM, supports rollback) |
| Manual refinement command | Optional `/refine` slash command, registered through the host `commands` capability (`@deepseek-ai/dsh-commands`) when present |
| Memory lifecycle | Manual archive/unarchive/pin through refinement metadata; archived entries are hidden from injection and skill materialization |
| Ranked injection | Queries the latest effective direct-user message (up to 400 chars), ranks title matches above content matches, then applies freshness/id tie-breaks and a per-kind cap |
| Session wrap-up | Optional `harness_wrapup` tool gives mechanical keep/promote/archive advice; promotion is copy-only and conflicts return a deterministic error |
| In-session review trajectory | Rebuilt from session logs (tail-biased truncation) |
| Invariant guard | `harness/refinement` event validation + batched failure reporting |
| Explicit A/B benchmark | Single `harness_benchmark` action tool: fixed frozen cases, pre-refinement reference snapshots, and same-round reference/candidate A/B runs with code-owned decisions |

## Architecture

```
src/
  domain.ts      event declaration merging (SessionEventMap / MessageSourceMap / cordis Events)
  types.ts       HarnessState / RefinementProposal / RefinementResult and other types
  storage.ts     disk read/write of state and history (atomic writes, corruption degradation, local/global merge, jsonl history)
  refine.ts      validation, application, rollback (baseline conflict detection, version increments, growth limit)
  skills.ts      SKILL.md rendering + file reconciliation (generated skills are real dsh skills)
  render.ts      model-facing overview / summary / history rendering (ranked injection)
  usage.ts       injection telemetry keys and in-memory usage aggregation
  wrapup.ts      deterministic session wrap-up suggestions (keep/promote/archive)
  planner.ts     LLM planning prompts and JSON parsing (plan / auto-refine review prompts)
  store.ts       HarnessStore: combined storage + event publishing (session events + agent-scoped events)
  complete.ts    completeViaAgent: completion through ctx.get('llm')
  benchmark.ts   benchmark cases/snapshots + atomic benchmark store persistence
  evaluate.ts    isolated per-cell executor/reviewer evaluation (evidence + score)
  score.ts       code-owned aggregation and ACCEPTED/REJECTED decisions
  tool.ts        harness_refine / harness_wrapup / harness_benchmark tools
  projection.ts  pre-step projection (digest dedup, <harness_state> injection)
  driver.ts      automatic refinement driver (turn-interval gate / compaction gate / cooldown / re-entry guard)
  invariant.ts   runtime invariant plugin
  index.ts       plugin entry and Config
tests/           23 test files, 287 cases (storage / store / refine / rules / planner / driver / approval / audit / logfile / skills / invariant / plugin integration / rank / projection / archive / usage / wrapup / benchmark / evaluate / score / isolation / tool / benchmark integration)
```

### Data layout

```
<harnessRoot>/                      shared ESP experience root; defaults to ~/.dsh/harness/
  harness_state.json                cross-session global state (ESP)
  refinements.jsonl                 global refinement history (append-only, ESP)
  reviews.jsonl                     cross-batch gate/audit history (ESP extension)
  continual-harness.log             continual-harness implementation log (JSONL, 0600)
  continual-harness.log.1           rotated continual-harness log
  usage.events.jsonl                append-only injection telemetry (lazily loaded into memory on first access)
  benchmark/                        explicit benchmark store (validation layer)
    cases.json                      fixed benchmark cases (draft/frozen + frozen material hashes)
    snapshots/<snapshotId>.json     captured reference snapshots (read-only merged harness state)
    runs.jsonl                      append-only A/B run records (cells + evidence + code-owned decision)
  sessions/<sessionKey>/
    harness_state.json              session-local state (shadows same-id global entries)
    refinements.jsonl               session refinement history
```

- **Skills are real dsh skills:** applied skill edits materialize as `<name>/SKILL.md` bundles (with provenance metadata) under `Config.skillsDir`, kept in sync by deletes/rollbacks without touching user-owned skills in the same directory.

### Experience Solidification Protocol (ESP)

The Experience Solidification Protocol (ESP) is the **protocol surface** of this capability set, decoupled from this package's implementation:

| Protocol element | Carrier | Description |
| --- | --- | --- |
| Experience state schema | `harness_state.json` (`schemaVersion: 1`) | Four kinds of entries — `prompt / memory / skill / subagent` — each with `id / kind / version / content / updatedAt` |
| Experience history | `refinements.jsonl` (append-only) | One `RefinementResult` record per apply/rollback; rollback by id |
| Refinement event | session event `harness/refinement` (retired) | Written on apply/rollback by builds up to 0.3.0; this build never appends it and keeps only its payload type declared for legacy compatibility |
| Refinement notification | agent event `harness/refined` | Payload `{agent, result}`; subscribable by invariant and other plugins |
| Experience injection | message source `plugin` (`form: instructions`, `digest` in the content marker) | Pre-injected into the model context; deduplicated by digest change. The retired `harness-state` kind is still recognized so old logs replace their block instead of duplicating it |

Any dsh plugin can read and write experience through this protocol (write state files, append history, publish events, inject messages); this package is the protocol's **reference implementation and primary consumer** (planning / refinement / projection / automatic gate).

## Mounting (dsh profile)

Install into a profile in one line (published to npm):

```sh
dsh plugin --profile <name> add dsh-continual-harness
```

The package declares `dsh.bundle`, so `dsh plugin` installs it as a profile
layer and applies its `cordis.patch.yml`. Update with
`dsh plugin --profile <name> update dsh-continual-harness@latest`.

Manual overlay (before publish, or to pin a local checkout): apply
[cordis.patch.yml](cordis.patch.yml) onto the profile, e.g.
`~/.dsh/profiles/<name>/cordis.patch.yml`; a patch layer must be a
**top-level YAML array** (`insert` rows append plugin entries; id-targeted rows override an existing row):

```yaml
- insert:
    - id: continual-harness
      name: dsh-continual-harness
      config:
        defaultGlobal: true
```

Prerequisites: the `tools`, `agents`, `session`, `llm`, `systemPrompt` capability plugins must load before this plugin (its `inject` declaration enforces that; mounting is deferred until they load).

### dsh version compatibility

Verified against dsh `0.1.2-alpha.3`; peer floors stay `>=0.1.0-rc.6`, so older dsh releases keep working. Since dsh `0.1.2-alpha.3` no longer provides `@deepseek-ai/dsh-home-paths` inside the profile bundle, the plugin declares it as a hard dependency; `@deepseek-ai/dsh-invariants` is used for types only (dev-time) and is not required at runtime.

## Config

| Field | Default | Description |
| --- | --- | --- |
| `harnessRoot` | dsh data dir `harness/` | State root directory (temporary dir in tests) |
| `skillsDir` | `$DSH_HOME/skills` | Directory where skill entries materialize as dsh SKILL.md bundles (dsh's user skill root) |
| `defaultGlobal` | required | Target scope when the tool call omits `global` |
| `maxTrajectoryChars` | 12000 | Max characters of the planning trajectory (two-layer signal + digest summary; `plannerPrefixCache`-route dependent) |
| `plannerMaxTokens` | 32000 | Max tokens for the planner LLM call |
| `plannerPrefixCache` | `auto` | Planning input route: `auto` (Route A warm session prefix when the session shows `cacheReadTokens > 0`, falling back to Route B on a truncated reply), `session` (always Route A), `off` (always Route B summary) |
| `plannerPrefixMaxChars` | 12000 | Tail-biased character cap for the Route A session prefix (`deriveMessages` text) |
| `trajectorySignalRatio` | 0.5 | Fraction of the Route B trajectory budget kept verbatim (signal layer) vs digested |
| `autoRefine` | `{turnInterval: 25, compact: true, cooldownMs: 1200000}` | Auto-refine: turn-interval gate, compaction-end gate, cooldown, disable switch |
| `requireGlobalApproval` | `false` | Require explicit human approval before a global write commits (conservative mode) |
| `maxInjectedEntriesPerKind` | `6` | Positive-integer cap (step 1, minimum 1) for ranked injected entries per kind |
| `wrapupEnabled` | `true` | Register the optional `harness_wrapup` session wrap-up tool |
| `diagnosticsEnabled` | `true` | Run post-apply structural diagnostics after each committed refinement |
| `securityEnabled` | `false` | Enable the local security (credential-pattern) diagnostic provider |
| `auditReviews` | `true` | Append every gate verdict to `reviews.jsonl` under the harness root |
| `logToFile` | `true` | Persist harness logs to `continual-harness.log` (JSONL, `0600`, rotated) |
| `logMaxBytes` | `5242880` (5 MB) | Rotation cap for the harness log file |
| `maxEntryGrowth` | `0.5` | Per-commit entry growth fraction cap; `0` disables the check |
| `protectedKinds` | `['skill']` | Kinds the automatic path may not modify (reserved; per-entry `protection` is the enforced guard) |
| `benchmark` | `{enabled: true, defaultRuns: 1, maxRuns: 3, passThreshold: 60, regressionTolerance: 0, maxFailedCells: 0}` | Explicit `harness_benchmark` tool: iterations per case per side, run cap, report-only pass line, non-regression tolerance, max failed candidate cells |

## Refining

Two entry points: the `harness_refine` tool (LLM-callable) and the `/refine` slash command (when the host provides a `commands` capability).

**`harness_refine`** — `mode: 'plan'` (default) plans from instructions and commits atomically; `mode: 'rollback'` takes a `rollbackId` plus an explicit `--local` / `--global` scope to revert a committed refinement. Global writes require human approval when `requireGlobalApproval` is `true`.

**`/refine`** — same semantics, human-typed:

```sh
/refine --local organize my memories
/refine --global <instructions>
/refine rollback <id> --local
/refine rollback <id> --global
```

Bare `/refine` plans with no instructions in the default scope. Output: `status`, `scope`, `refinement`, `applied`, `rejected`, `summary`, plus a `diagnostics:` line when enabled.

## Governance

Every write path funnels through three guardrails: **impact minimization** (fixed contract validation; `update`/`delete` require a one-line `reason`; `maxEntryGrowth` caps per-commit growth), **legality hard rejects** (`base_system_prompt` and protected entries are immutable; global entries are read-only during a `local` refinement), and a **necessity soft gate** (a declined review never reaches the store). Every committed refinement rolls back by id.

Global writes are **zero-approval by default**; set `requireGlobalApproval: true` to ask the user first. Watch the plugin log live with:

```sh
tail -f ~/.dsh/harness/continual-harness.log
```

## Benchmark

The validation layer is **explicit and single-entry**: one `harness_benchmark` action tool drives the whole workflow and never auto-triggers a refinement — nothing in the benchmark path starts a `harness_refine` or the automatic gate, and a `REJECTED` decision is reported and recorded only, never rolled back. The store lives under `<harnessRoot>/benchmark/` (see the data layout above).

The minimal sequence is `new → add-case → freeze → capture-reference → apply refinement → run → status` (frozen case material is immutable and hashed; `status` lists cases, snapshots, and recent runs). Two steps carry real subtleties:

- `capture-reference` must run **BEFORE** the refinement you want to validate: the candidate is later derived as *the captured reference plus exactly that refinement*, so capturing after the change would make the delta unprovable.
- `run` evaluates the named refinement A/B against the reference (`reference_snapshot_id` + `refinement_id`). The candidate must be the **single specified delta** — derived from the reference plus the refinement's recorded applied edits and proved in code before any evaluation; a drifted or multi-change candidate is refused (`benchmark:run:candidate-delta`). Both sides run the same frozen cases in stored order with the same `runs`/`provider`/`model`.

A `run` returns the code-owned decision (`src/score.ts`), not a model verdict:

```json
{
  "action": "run",
  "ok": true,
  "run_id": "run-...",
  "refinement_id": "refine-1",
  "status": "ACCEPTED",
  "reference_overall": 70,
  "candidate_overall": 90,
  "regression_cases": [],
  "failed_cells": 0,
  "feedback": ["reference ok", "candidate better"],
  "auto_rollback": false,
  "runs": 1,
  "cells": 2
}
```

- Scores are `0..100` per cell; a failed cell carries `score: null` — failure is never counted as `0` — and is excluded from the overall means.
- `passThreshold` (default `60`) is **report-only**: it never gates acceptance. A run is `ACCEPTED` only when neither side lacks usable cells, candidate failed cells stay within `maxFailedCells`, and no overall or per-case regression exceeds `regressionTolerance` (default `0`).
- Every run appends its full record (cells with executor evidence + the decision) to `benchmark/runs.jsonl`; evaluation reads only the captured snapshots and writes only that record, never touching `reviews.jsonl`, the harness state, injection telemetry, or skill files.

## Development

The plugin is self-contained: `devDependencies` pin the published
`@deepseek-ai/*` packages (rc versions), so `pnpm install`, `pnpm run
typecheck`, `pnpm test`, and `pnpm run build` (tsc emits
`lib/types/*.js + *.d.ts`; the `"."` and `"./invariant"` exports point at the
artifacts) all work in a clean checkout — CI and the OIDC release workflow
run the same steps. `peerDependencies` declare the semver ranges consumers
(host dsh installations) must satisfy.

Plugin builds up to 0.3.0 logged the injected overview under a plugin-defined
`harness-state` message source. The released Session format migrations only
classify platform source kinds, so one such message makes the whole stored
artifact unreadable (`cannot safely transform unclassified message source`)
once a host reads it with a v3-capable dsh. This build logs a classified
`plugin` source instead; stored logs of any generation are repaired offline
with `node scripts/repair-harness-state-logs.mjs` (dry run by default; `--apply`
backs each artifact up and replaces it atomically, and artifacts written within
`--min-age-seconds` are skipped — see `--help`).

## Known Limitations and Deferred Work

- No end-to-end tests with a real LLM: `completeViaAgent` depends on the loaded `llm` capability and provider/model configuration; tests cover the planning/review paths with a stub `Complete`. Real e2e requires `DEEPSEEK_API_KEY`.
- `compaction/end` is not part of the plugin's type union; the driver triggers it via string comparison after type narrowing, and the gate is silently skipped when the compaction capability is not loaded.
- Projection dedup is an in-process `WeakMap<Agent, digest>`: the first step after a session restart re-injects (stateless and idempotent, but one extra injection).
- Concurrent writes are last-writer-wins: multiple processes refining the same directory concurrently may overwrite each other; baseline conflict detection during planning can only catch read-after-write races, not serialize them.
- A failed automatic refinement degrades silently (only logged) and never interrupts the session.
- A content-shrink guard (rejecting updates that shrink an entry too far in one commit) is a planned follow-up and is not yet implemented; today only `maxEntryGrowth` caps how much an update may grow an entry.
- A dedicated governance tool entry is deferred.

Install

dsh plugin --profile web add github:jasen215/dsh-continual-harness

Profile: web

  • This package builds from source on install. pnpm will ask you to allow its build script — that is permission to run the package’s code on your machine, outside the agent sandbox. Only allow sources you trust.
  • This source has no pinned commit, so a later push upstream changes what installs. Prefer pinning a commit.
Source