Bundle
dsh-kokoro
DeepSeek Harness plugin: speak text with a local Kokoro TTS HTTP server and return the wav path, as the kokoro_tts tool. Ships no weights.
- Source
- jwilson411
- License
- MIT
- Updated
- Updated 2 days ago
Readme
# dsh-kokoro
A small **DeepSeek Harness function plugin** that speaks text out loud. It
registers exactly one model-facing tool — `kokoro_tts` — which POSTs text to a
local [Kokoro TTS server](https://github.com/jwilson411/kokoro-tts-api), writes
the wav that comes back, and reports the path.
**No weights ship here.** No ONNX, no `.pt`, no voice bins, no embeddings —
nothing in this repository is model data, and nothing is downloaded at install
time. The model lives behind the server; this package is the wire to it. What
it does contain is a list of nineteen voice *names* and two numeric bounds,
mirrored from that server so a bad argument fails here instead of costing a
round trip.
It is also not an audio toolkit. There is no playback, no format conversion, no
mixing, no streaming, no caching, no transcription. It answers one question —
*what would this text sound like?* — and hands back a file path.
No API key, no credentials, no state.
## Install
```sh
dsh plugin --profile web add github:jwilson411/dsh-kokoro
```
`dsh plugin` forwards to pnpm inside `$DSH_HOME/profiles/web`, then reconciles
the profile against the installed state: because this package's manifest
declares `dsh.bundle.patch`, it is appended to the profile manifest's ordered
`dsh.profile.bundles` list and its patch becomes a layer.
Remove it the same way, with `remove` in place of `add`.
## Point it at a server
The plugin's built-in default base URL is **`http://127.0.0.1:8000`**.
The companion server, [`jwilson411/kokoro-tts-api`][api], listens on **8765** by
default (`KOKORO_PORT`). So if that is the server you are running — and it
probably is — set the base URL to `http://127.0.0.1:8765`. Either the patch row:
```yaml
- id: kokoro
config:
baseURL: http://127.0.0.1:8765
```
or the environment:
```sh
export DSH_KOKORO_BASE_URL=http://127.0.0.1:8765
```
The order is most specific first: the plugin's patch row, then
`DSH_KOKORO_BASE_URL`, then the default. A row that names `baseURL` is taken at
its word even when blank — writing the key and leaving it empty is a
misconfiguration, and it fails at load rather than quietly talking to a server
you did not name. Whatever wins is normalized (trailing slash stripped) and
scheme-checked immediately: anything that is not `http:` or `https:` is refused,
so a `file:` base cannot turn the one outbound request into a local read.
Bring the server up first — see its README for the weights step, which happens
over there, not here:
```sh
KOKORO_PORT=8765 make run # in a checkout of jwilson411/kokoro-tts-api
curl -s http://127.0.0.1:8765/health
```
## Pinned DSH release candidate
This package is written and tested against the pinned release candidate
**`0.1.1-rc.2`** — the current `@deepseek-ai/dsh` release and the matching
`@deepseek-ai/dsh-tools@0.1.1-rc.2`, which is pinned exactly in
`devDependencies` so tests run against one known API. The peer range is
`^0.1.1-rc.2`, matching how the harness's own tool packages declare it.
Note that `@deepseek-ai/dsh-tools`'s npm `latest` tag still points at the older
`0.0.1-rc.1`; the `0.1.1-rc.2` line is published under `next`. Pin explicitly
rather than relying on the tag.
There are no runtime dependencies.
## What it registers
| | |
|---|---|
| Cordis plugin id | `kokoro` (the row id in `cordis.patch.yml`) |
| Injects | `tools` — a hard dependency; the plugin waits rather than degrading |
| Tool | Arguments | Returns |
|---|---|---|
| `kokoro_tts` | `text` (string, required, 1–8000 chars), `voice` (string, optional, allowlisted, default `am_michael`), `speed` (number, optional, 0.5–2.0, default 1.0) | `{ path, voice, speed, bytes, text_chars, plugin }` |
```json
{
"path": "/tmp/dsh-kokoro/kokoro-am_michael-0f9c1f2e-....wav",
"voice": "am_michael",
"speed": 1,
"bytes": 122444,
"text_chars": 41,
"plugin": "dsh-kokoro"
}
```
Nothing is played automatically. The path is reported so a caller can open it.
### Voices
Nineteen, allowlisted. `am_`/`af_` are American male and female, `bm_`/`bf_`
British male and female:
`am_michael` (default), `am_adam`, `am_liam`, `am_eric`, `am_james`,
`am_william`, `am_caleb`, `am_david`, `am_ethan`, `bm_daniel`, `bm_george`,
`bm_lewis`, `bm_oliver`, `af_heart`, `af_nova`, `af_sarah`, `af_bella`,
`bf_emma`, `bf_isabella`.
Any other name is refused with `KOKORO_BAD_VOICE` before the socket opens.
### Where the wav goes — and why you cannot choose
The tool has **no `path`, `output`, or `filename` argument**, and will not grow
one. A model choosing where bytes land is a file-write primitive wearing a
text-to-speech costume; a `path` argument on this tool would let any prompt that
reaches it write attacker-shaped bytes to an attacker-chosen location. So the
plugin picks: a file named `kokoro-<voice>-<uuid>.wav` inside a directory it
owns — `os.tmpdir()/dsh-kokoro` by default, created mode 0700, files written
0600 — and tells you the path afterwards. An operator can move that directory
with `config.outputDir` in the patch row. A caller cannot.
Nothing is cleaned up on your behalf; these are files under your temp directory,
and their lifetime is your system's business.
## Talking to Kokoro
Exactly one request is ever made, to exactly one URL: `POST {baseURL}/tts`, with
`{text, voice, speed}` as JSON, expecting `audio/wav` back. Redirects are **not**
followed — a 3xx is reported as `KOKORO_HTTP_ERROR` rather than carrying the
text being spoken to whatever host the `Location` header names.
Every request is bounded twice and fails closed on either bound: an
`AbortSignal` deadline (30s — synthesis is not instant) and a response byte cap
(10 MiB) enforced while the body streams, so an oversized response is abandoned
rather than buffered. Both are configurable from the patch row:
```yaml
- insert:
- id: kokoro
name: dsh-kokoro
config:
baseURL: http://127.0.0.1:8765
timeoutMs: 30000
maxBytes: 10485760
outputDir: /var/tmp/dsh-kokoro
```
An id-targeted patch replaces the row's whole `config` rather than merging into
it, so an override must restate the fields it keeps.
What comes back is checked before any of it touches disk: the content type must
be `audio/*`, and the first twelve bytes must be a RIFF container declaring
WAVE. A proxy login page or a JSON error body served with a 200 fails as
`KOKORO_BAD_AUDIO` instead of landing on disk with a `.wav` name.
Failures carry a stable `code` — `KOKORO_BAD_TEXT`, `KOKORO_BAD_VOICE`,
`KOKORO_BAD_SPEED`, `KOKORO_BAD_BASE_URL`, `KOKORO_HTTP_ERROR`,
`KOKORO_BAD_AUDIO`, `KOKORO_RESPONSE_TOO_LARGE`, `KOKORO_TIMEOUT`,
`KOKORO_UNREACHABLE`, `KOKORO_WRITE_FAILED` — so a caller can tell a dead port
from a rejected voice without matching on prose. Requests identify themselves
with a `User-Agent` naming the plugin and this repository.
## Headless use
The tool factory takes both seams — the `fetch` to use and the directory to
write into — so you can drive it from a plain Node script without booting a
profile:
```js
// say.mjs — node say.mjs
import { createKokoroTtsTool } from 'dsh-kokoro'
// Omit `fetch` to use the global one and talk to a real server.
const tts = createKokoroTtsTool({ baseURL: 'http://127.0.0.1:8765' })
const { path, bytes } = await tts.execute(
{ text: 'The kettle is on.', voice: 'bm_george' },
{ signal: AbortSignal.timeout(60_000) },
)
console.log(`${bytes} bytes at ${path}`)
```
### `KOKORO_STUB=1`
With `KOKORO_STUB=1` in the environment **and** no injected `fetch`, no request
is made at all: a short silent wav is generated locally so the rest of the
pipeline — validate, write, report a path — can be walked with no server
running. It is silence, not speech; nothing here synthesizes audio and no model
is involved. Treat it as a demo and smoke-test hatch, not a fallback: it never
fires when a `fetch` was supplied, and it never masks a real failure.
## Tests
```sh
npm install
npm test
```
The suite is offline and needs no Kokoro server. Every request is answered by an
injected `fetch` double, and each test file replaces `globalThis.fetch` with a
guard that throws, so a code path that reached for the real network fails as a
test failure rather than a live request. Files are written into a throwaway
directory under `os.tmpdir()`, one per test, removed when it ends. CI runs it on
Node 22.x and 24.x with `contents: read` and no secrets.
## Licence
MIT — see [LICENSE](LICENSE). Copyright (c) 2026 Justin Wilson.
The Kokoro-82M model and its weights are not part of this package and are not
covered by this licence; see [`jwilson411/kokoro-tts-api`][api] for those.
[api]: https://github.com/jwilson411/kokoro-tts-api
Install
dsh plugin --profile web add github:jwilson411/dsh-kokoro
Profile: web
With the hub plugin installed, ask your agent to install it by name — it resolves the same plan shown here.
dsh plugin --profile web add github:stvlynn/dsh.fish#path:packages/dsh-plugin-hub
install dsh-kokoro from the hub
- This source has no pinned commit, so a later push upstream changes what installs. Prefer pinning a commit.