Bundle
dsh-llm-gate
Per-provider concurrency gate for DeepSeek Harness LLM requests
- Source
- d3vmeh
- stars
- 1 stars
- License
- MIT
- Updated
- Updated 4 days ago
Readme
# dsh-llm-gate
Per-provider concurrency gate for [DeepSeek Harness](https://github.com/deepseek-ai/deepseek-harness) model requests.
If a provider can only serve a fixed number of requests at once (e.g a local `llama-server` with `--parallel 1`), every extra request is deferred by the server with nothing sent back. The client cannot tell "waiting for a slot" from "dead", and Node HTTP layer times out after 300 seconds with `terminated`. In practice this happens when there is overlap between a subagent and the main agent or compaction and the agent.
This plugin holds surplus requests inside dsh instead. A request waits in a FIFO queue before any HTTP request is made so no timeout is running while it waits. When a slot frees, the next request is dispatched.
## Install
```
dsh plugin --profile web add dsh-llm-gate
```
Then configure the providers to gate in `~/.dsh/profiles/web/cordis.patch.yml`:
```yaml
- id: llm-gate
config:
providers:
llamacpp:
maxConcurrent: 1
maxQueued: 16
queueTimeoutMs: 3600000
```
The provider key is the route name from your `llm-pi-ai.providers` (or other adapter) settings. Providers not listed are not gated. Restart `dsh web` and open a new session.
Check the composed config with `dsh --profile web --dump-config`.
## Settings
| Setting | Required | Meaning |
|---|---|---|
| `maxConcurrent` | yes | Requests allowed in flight to this provider. For llama.cpp, match `--parallel`. |
| `maxQueued` | no | Requests allowed to wait. Beyond this, a request fails at once with `QUEUE_FULL`. Default: unlimited. |
| `queueTimeoutMs` | no | Longest a request may wait for a slot before failing with `QUEUE_TIMEOUT`. Default: wait indefinitely. |
Queue failures end the turn with the code shown. They are not retried by `dsh-llm-retry`.
## What you will see
The plugin prints a line to the dsh terminal only when a request has to wait:
```
llm-gate: llamacpp session=a61e6e40 queued (depth 1)
llm-gate: llamacpp session=a61e6e40 dispatched after 5730ms
```
`purpose=compaction` or `purpose=session-title` is added for auxiliary requests. Requests that get a slot immediately print nothing.
## Notes
- This gate serializes requests so it does not make a single-slot server faster. For parallelizing, give llama.cpp more slots (`--parallel 2 --kv-unified`) and raise `maxConcurrent` to match.
- Waiting time is not counted by the adapter's `streamIdleTimeoutMs` because the adapter is not called until the slot is acquired. You still need `streamIdleTimeoutMs` large enough for your prompt processing time (see the `llm-pi-ai` provider settings).
- A queued request is cancelled through its abort signal. Dropping the stream without aborting leaves the request queued until a slot frees, at which point it dispatches and is closed immediately.
- Requires the `llm` service; hooks the `llm/stream` waterfall, so it covers every model request in the host: agents, subagents, compaction, and title generation.
## License
MIT
Install
dsh plugin --profile web add github:d3vmeh/dsh-llm-gate
Profile: web
With the hub plugin installed, ask your agent to install it by name — it resolves the same plan shown here.
dsh plugin --profile web add github:stvlynn/dsh.fish#path:packages/dsh-plugin-hub
install dsh-llm-gate from the hub
- This source has no pinned commit, so a later push upstream changes what installs. Prefer pinning a commit.