Bundle
@dsh-extension/dsh-generation-image
On-demand image generation for DeepSeek Harness (DSH): a generate_image tool that calls your own OpenAI-compatible image URL + API key and delivers the image into the session
- Source
- sfyyy
- stars
- 1 stars
- License
- MIT
- Updated
- Updated 13 days ago
Readme
# @dsh-extension/dsh-generation-image
> On-demand image generation for DeepSeek Harness (DSH) sessions.
[](https://developer.mozilla.org/en-US/docs/Web/JavaScript)
[](https://www.npmjs.com/package/@dsh-extension/dsh-generation-image)
[](./LICENSE)
[](https://www.npmjs.com/package/@deepseek-ai/dsh)
[中文文档](./README.zh-CN.md)
A DSH plugin that gives a **DeepSeek session on-demand image generation**:
the agent gets a `generate_image` tool that calls **your own** OpenAI-compatible
image URL + API key (`POST /images/generations` or `/images/edits`) and delivers
the generated image into the session.
- **Text-to-image and image-to-image** — generate from a prompt, or edit and
combine one or more DSH image attachments through an OpenAI-compatible
`/images/edits` endpoint.
- **Bring your own image endpoint** — OpenAI-compatible image services such as
OpenAI, xiaoyaoapi, vLLM image models, or local gateways.
- **The image lands in the session UI** — generated bytes are durably saved as
DSH attachments, surfaced as an assistant-side image message, and rendered as a
clickable **thumbnail** in the conversation. Click it to **enlarge**, then
**download or close** from the enlarged view.
- **The text model stays safe** — DeepSeek is a text-only model, so image blocks
are rewritten to text markers for every text-model request (same mechanism as
[dsh-vision-bridge](https://github.com/sfyyy/dsh-vision-bridge)); the session
log and UI keep the real image.
- **Mirrors the proven GPT2Image request shape** — `stream: true`,
`partial_images: 1`, SSE `image_generation.completed` events (with a plain
JSON fallback), the exact endpoint used by the GPT2Image app
(`https://api.xiaoyaoapi.cc/v1`, model `gpt-image-2`).
- **b64_json and url results both supported** — channels that return the image
inline (`b64_json`, SSE or plain JSON) work as before; channels that return a
remote `url` (e.g. `data[].url` or url-carrying SSE events) are handled
transparently by downloading the image bytes before saving the attachment.
## How it works
```text
agent calls generate_image(prompt, size?, quality?, count?, referenceImageIds?)
│
▼
referenceImageIds omitted → POST {baseUrl}/images/generations (JSON)
referenceImageIds provided → attachments.readImage() → POST {baseUrl}/images/edits
(multipart, repeated image[] fields)
│
▼
SSE (image_generation.partial_image / .completed → b64_json or url)
or plain JSON (data[].b64_json / data[].url)
│ (url results are downloaded automatically)
▼
magic-byte sniff → attachments.saveImage() → durable image ref
│
▼
tool result: text envelope + image block(s), plus an assistant-side display
message → the conversation renders a clickable thumbnail on the left
│
▼
thumbnail click → built-in lightbox (enlarge) with 下载原图 (download) + close
│
▼
deriveMessages() rewrites image blocks to text markers (text model never
sees image blocks); llm.resolveModelInfo admission bypass lets the messages
enter the agent.
```
## Viewing generated images
The DSH Web UI renders every generated image as a **thumbnail** in the
conversation (an assistant-side image message). Multiple images in one message
are **stacked vertically**; click any image to open the enlarged preview:
- **下载原图** — the download button in the top-right of the enlarged preview
saves the current full-resolution image to disk.
- **关闭** — the close button in the top-right (or clicking the mask).
- **左右切换** — with multiple images, the preview shows `n / N` and provides
previous/next buttons on the sides.
No extra configuration is needed; the stacked display and preview controls are
enabled as soon as the plugin's client bundle is loaded (refresh the DSH web
page after upgrading the plugin).
## Installation
Install from the npm registry (not a local checkout) — one command:
```sh
# if you already have the `dsh` CLI on PATH:
dsh plugin --profile web add @dsh-extension/dsh-generation-image
# or, if you have been using npx all along:
npx @deepseek-ai/dsh@0.1.0-rc.6 plugin --profile web add @dsh-extension/dsh-generation-image
```
> The `--profile` flag targets the profile you boot (`web` is the browser UI
> profile). Omit it or adapt it if your profile has a different name.
>
> After a new client bundle is added, restart `dsh web` once so the UI picks it up.
### Local development (like vision-bridge)
From a local checkout, link it into your profile like the sibling plugins, then
restart `dsh web`:
```jsonc
// ~/.dsh/profiles/web/package.json → dependencies
{
"@dsh-extension/dsh-generation-image": "link:/path/to/dsh-generation-image"
}
// ~/.dsh/profiles/web/package.json → dsh.profile.bundles
"@dsh-extension/dsh-generation-image"
```
or, for a running instance, use the DSH super-injector:
```sh
dsh plugin inject /path/to/dsh-generation-image
```
## Configuration
Configure it in **Settings → Generation Image** (DSH Web), or edit
`~/.dsh/generation-image.json`. **The plugin ships with empty credentials by
default** — you supply your own endpoint URL + API key (via the Settings page,
env vars, or the config file):
```json
{
"enabled": true,
"baseUrl": "https://your-image-endpoint.example/v1",
"apiKey": "sk-xxxx",
"model": "gpt-image-2",
"size": "",
"quality": "auto"
}
```
- `baseUrl` — OpenAI-compatible image API root (`.../v1`); **empty by default**,
the plugin normalizes it and calls `${baseUrl}/images/generations` or
`${baseUrl}/images/edits` according to the tool arguments.
- `apiKey` / `apiKeyEnv` — **empty by default**, mutually exclusive. A directly
entered key is synced to the DSH credential store and referenced as
`DSH_GENERATION_IMAGE_API_KEY`.
- `model` — the image model id (default `gpt-image-2`).
- `size` — default size hint, **empty by default = unrestricted**: the model
passes any size per call, or `"auto"` to let the API decide. No value is
hardcoded or restricted.
- `quality` — default quality hint: `auto` (default; omitted from the request,
the API decides) or any value the model/endpoint accepts.
- `enabled: false` disables the whole chain: no tool registration, no image
rewriting, no admission bypass (native behavior restored).
**Precedence (highest wins):** Settings page (with schema defaults) →
environment variables → config file.
**Environment overrides:** `DSH_GENERATION_IMAGE_BASE_URL`,
`DSH_GENERATION_IMAGE_API_KEY`, `DSH_GENERATION_IMAGE_API_KEY_ENV`,
`DSH_GENERATION_IMAGE_MODEL`, `DSH_GENERATION_IMAGE_SIZE`,
`DSH_GENERATION_IMAGE_QUALITY`, `DSH_GENERATION_IMAGE_ENABLED`.
## `generate_image` tool
- **Arguments**
- `prompt` (required): a detailed description of the image to generate;
- `size` (optional, **unrestricted**): pass any size the endpoint accepts
(e.g. `1024x1024`, `1024x1792`, `1792x1024`), or `"auto"`/omit to let the
API decide. For true 4K use **`3840x2160`** or **`3840x3840`** — the current
upstream max edge is 3840px; `4096` exceeds the limit and fails. The
plugin's bundle patch also raises DSH attachment-local limits
(`normalizedImageMaxDimension: 4096`, `normalizedImageMaxBytes: 26214400`)
so saved 4K images are not compressed down to 2048;
- `quality` (optional, **unrestricted**): common values are `auto` (default),
`low`, `medium`, `high`, or any value your endpoint accepts;
- `count` (optional, 1–4): how many images to generate (default 1). The
upstream endpoint currently requires `n=1`, so the plugin automatically
splits `count > 1` into multiple `n=1` requests. The DSH client renders
multiple images in one assistant message stacked vertically.
- `referenceImageIds` (optional): ordered, unique image attachment ids from
the current conversation. Omit for text-to-image; provide one or more for
image-to-image editing or combining references. The deployment's DSH image
count limit applies.
- **Behavior**: calls the configured image endpoint → parses the SSE stream
(or plain JSON) → sniffs the real media type from magic bytes → durably saves
each image through the DSH attachment service → returns a text envelope plus
one image block per generated image.
- **Result**: the generated image(s) appear in the session log and Web UI. The
text model receives a text marker instead of the image block.
- **Image-to-image usage**: upload one or more images and ask to modify,
restyle, or combine them. The bundled Skill extracts their attachment ids and
passes `referenceImageIds`. If no reference exists, the agent asks for an
upload instead of silently falling back to text-to-image.
## Verify
```sh
npm test
```
The suite covers: tool registration gating, the image-API call (SSE + plain
JSON), multi-reference `/images/edits` multipart requests, attachment isolation,
canonical value + rendered output, nested `run_code` behavior,
image-marker rewriting (session log untouched), admission bypass on/off,
tool visibility after downstream filters, disabled behavior, and config/env
precedence.
## Development
From a local checkout:
```sh
dsh plugin inject /path/to/dsh-generation-image
```
## Search keywords
`deepseek` · `deepseek-harness` · `dsh` · `plugin` · `image generation` ·
`text-to-image` · `image-to-image` · `img2img` · `generate image` · `gpt-image` · `OpenAI-compatible` ·
`images API` · `xiaoyaoapi` · `LLM agent`
## License
[MIT](./LICENSE)
Install
dsh plugin --profile web add github:sfyyy/dsh-generation-image
Profile: web
With the hub plugin installed, ask your agent to install it by name — it resolves the same plan shown here.
dsh plugin --profile web add github:stvlynn/dsh.fish#path:packages/dsh-plugin-hub
install dsh-extension-dsh-generation-image from the hub
- This source has no pinned commit, so a later push upstream changes what installs. Prefer pinning a commit.