Skip to content
dsh.fish
Bundle

dsh-skill-eval

Skill-trigger evaluation: an LLM judge recreates the DSH skill catalog and measures how reliably a skill description routes matching queries.

Source
renjianguojinqianfan
stars
2 stars
License
MIT
Updated
Updated 13 hours ago

Readme

# dsh-skill-eval

[![CI](https://github.com/renjianguojinqianfan/dsh-skill-eval/actions/workflows/ci.yml/badge.svg)](https://github.com/renjianguojinqianfan/dsh-skill-eval/actions/workflows/ci.yml)

Skill-trigger evaluation plugin for DeepSeek Harness (DSH).

An LLM judge recreates the exact DSH skill catalog prompt and decides, for each
test query, whether the target skill should be triggered. The plugin reports
accuracy, precision, recall, false-positive/negative rates, and a confusion
matrix — a reproducible measure of how reliably a skill description routes
matching queries (and how often it over- or under-triggers).

## Install

From the repo root:

```bash
dsh plugin --profile web add ./dsh-skill-eval
```

Then configure the judge model route in your profile/overlay `cordis.patch.yml`:

```yaml
- id: skill-eval
  config:
    provider: <provider-id>
    model: <model-name>
```

The provider must be registered in the DSH LLM runtime (the same one your
profile uses for chat). The plugin validates the route at startup and warns if
the provider is not yet registered.

## Usage

Slash command:

```
/skill-eval <skill-name> [test-file]
```

Model-callable tool:

```
run_skill_eval(skill_name="<skill-name>", test_file="examples/dsh-plugin-eval.json")
```

`test-file` is optional; it defaults to `examples/dsh-plugin-eval.json` inside
the plugin package. Relative paths resolve against the plugin package directory.

## Test-case format

A JSON array of `{ query, should_trigger }` objects:

```json
[
  { "query": "add a tool to the harness that persists across restarts", "should_trigger": true },
  { "query": "help me write a Python script for this CSV", "should_trigger": false }
]
```

`category` is optional and reserved for future use.

## How it works

1. Enumerate the session's model-invocable skills (`ctx.skills.snapshot`).
2. Recreate the official catalog message verbatim (`<system-reminder>` +
   `<available_skills>` + normalized/truncated/escaped descriptions).
3. For each query, ask the judge model whether the target skill should be
   triggered, forcing a one-line `YES`/`NO` answer.
4. Compare against the expected label and aggregate metrics.

The evaluation measures the *judge model's* routing accuracy for the given
skill description. Swap `provider`/`model` in the config to test other judges.

## Development and tests

```bash
npm run check     # syntax check for every JS file
npm test          # node:test, including official catalog fidelity and mock ctx tests
npm run smoke     # 51 pure-function smoke assertions
npm pack --dry-run  # published file list check
bash scripts/mount-smoke.sh  # real DSH mount smoke in a scratch home
```

The catalog fidelity fixture pins the official `dsh-tool-skill@0.1.0-rc.6`
template. After a DSH upgrade, refresh the fixture from a local official
install and review the diff:

```bash
node scripts/refresh-catalog-fixture.mjs <path-to-dsh-tool-skill/lib/index.js>
```

## Files

- `index.js` — plugin entry: registers the `run_skill_eval` tool and the
  `/skill-eval` command.
- `runner.js` — catalog reproduction, judge LLM call, verdict parsing.
- `catalog.js` — pure functions: catalog message rendering and verdict parsing.
- `llm-helpers.js` — dependency-free `BlockAssembler`, `createUserMessage`,
  `deepFreeze` (mirrors the official dsh-llm pattern).
- `parser.js` — test-case JSON loading and validation.
- `metrics.js` — confusion matrix, metrics, and markdown report formatting.
- `examples/` — default test cases.

Install

dsh plugin --profile web add github:renjianguojinqianfan/dsh-skill-eval

Profile: web

  • This source has no pinned commit, so a later push upstream changes what installs. Prefer pinning a commit.
Source