Skip to content
dsh.fish
Bundle

dsh-llm-gate

Per-provider concurrency gate for DeepSeek Harness LLM requests

Source
d3vmeh
stars
1 stars
License
MIT
Updated
Updated 4 days ago

Readme

# dsh-llm-gate

Per-provider concurrency gate for [DeepSeek Harness](https://github.com/deepseek-ai/deepseek-harness) model requests.

If a provider can only serve a fixed number of requests at once (e.g a local `llama-server` with `--parallel 1`), every extra request is deferred by the server with nothing sent back. The client cannot tell "waiting for a slot" from "dead", and Node HTTP layer times out after 300 seconds with `terminated`. In practice this happens when there is overlap between a subagent and the main agent or compaction and the agent.

This plugin holds surplus requests inside dsh instead. A request waits in a FIFO queue before any HTTP request is made so no timeout is running while it waits. When a slot frees, the next request is dispatched.

## Install

```
dsh plugin --profile web add dsh-llm-gate
```

Then configure the providers to gate in `~/.dsh/profiles/web/cordis.patch.yml`:

```yaml
- id: llm-gate
  config:
    providers:
      llamacpp:
        maxConcurrent: 1
        maxQueued: 16
        queueTimeoutMs: 3600000
```

The provider key is the route name from your `llm-pi-ai.providers` (or other adapter) settings. Providers not listed are not gated. Restart `dsh web` and open a new session.

Check the composed config with `dsh --profile web --dump-config`.

## Settings

| Setting | Required | Meaning |
|---|---|---|
| `maxConcurrent` | yes | Requests allowed in flight to this provider. For llama.cpp, match `--parallel`. |
| `maxQueued` | no | Requests allowed to wait. Beyond this, a request fails at once with `QUEUE_FULL`. Default: unlimited. |
| `queueTimeoutMs` | no | Longest a request may wait for a slot before failing with `QUEUE_TIMEOUT`. Default: wait indefinitely. |

Queue failures end the turn with the code shown. They are not retried by `dsh-llm-retry`.

## What you will see

The plugin prints a line to the dsh terminal only when a request has to wait:

```
llm-gate: llamacpp session=a61e6e40 queued (depth 1)
llm-gate: llamacpp session=a61e6e40 dispatched after 5730ms
```

`purpose=compaction` or `purpose=session-title` is added for auxiliary requests. Requests that get a slot immediately print nothing.

## Notes

- This gate serializes requests so it does not make a single-slot server faster. For parallelizing, give llama.cpp more slots (`--parallel 2 --kv-unified`) and raise `maxConcurrent` to match.
- Waiting time is not counted by the adapter's `streamIdleTimeoutMs` because the adapter is not called until the slot is acquired. You still need `streamIdleTimeoutMs` large enough for your prompt processing time (see the `llm-pi-ai` provider settings).
- A queued request is cancelled through its abort signal. Dropping the stream without aborting leaves the request queued until a slot frees, at which point it dispatches and is closed immediately.
- Requires the `llm` service; hooks the `llm/stream` waterfall, so it covers every model request in the host: agents, subagents, compaction, and title generation.

## License

MIT

Install

dsh plugin --profile web add github:d3vmeh/dsh-llm-gate

Profile: web

  • This source has no pinned commit, so a later push upstream changes what installs. Prefer pinning a commit.
Source