Skip to content
dsh.fish
Bundle

dsh-multimodal

Multimodal eyes and hands for DeepSeek Harness: vision transcription, OCR, and text-to-image via OpenAI-compatible backends, with an in-conversation generated-image card.

Source
MC5lan
stars
4 stars
License
MIT
Updated
Updated 7 days ago

Readme

# dsh-multimodal — multimodal eyes and hands for DeepSeek Harness

English | [简体中文](README.zh.md)

Give DeepSeek a pair of eyes and a paintbrush: paste a screenshot or image into the conversation and the vision provider **you configured** first transcribes it verbatim (error messages, code, UI text preserved), then **DeepSeek keeps solving your problem** — all in the same turn, no extra steps. When an illustration is needed, DeepSeek automatically calls the image backend and the generated pictures appear right in the conversation.

> **Blank slate by design**: this plugin ships **no built-in models, providers, or backends**. Vision endpoints, image backends, and models are all declared by *you* — plug in whatever API you already have (DeepSeek, Zhipu, Aliyun, SiliconFlow, ModelScope, Xfyun, Qianfan, local Ollama, …). Nothing is preloaded, nothing is assumed.

> **Compatibility**: built for [DeepSeek Harness](https://github.com/deepseek-ai/deepseek-harness) `0.1.0-rc.6` (Web and headless). See [CHANGELOG](CHANGELOG.md).

## Features

| Scenario | Behavior |
|---|---|
| Plain-text chat | Straight to the DeepSeek API (unchanged) |
| Image + question (e.g. error screenshot) | Your configured vision provider "looks" first → transcribes to text → **DeepSeek continues from the transcription** (fix code, explain, propose); hitting "stop" aborts the vision call immediately |
| Attaching an image | **No more "current model does not support images"** |
| User asks for an image | DeepSeek calls `generate_image` → the configured image backend produces pictures shown in the conversation; **backend failover** tries the next backend if the active one fails (AUTH/aborted skips failover — no wasted quota) |
| Any image API | A `custom` backend plugs any non-OpenAI/DashScope API in via a small adapter file — no plugin code changes |
| Image card | Dedicated `generate_image` card: thumbnail grid, click-to-zoom lightbox, one-click download, prompt & model metadata, **crop-to-ask** (drag a region + ask), **copy params** (reproducible JSON), **retry** button on failure (refine-aware) |
| Extracting text from an image | DeepSeek can call `extract_text` (OCR) → Markdown / plain text / JSON; decoupled from the watch route, usable in any session |
| **Paste-key auto-connect** (0.7+) | Paste any API key into the quick-add box → platform auto-detected (key fingerprint → `/models` probing) → endpoint + allow-list + credential + model list + feature enablement in one step |
| Vision platforms | `extraProviders` accepts **any OpenAI-compatible vision endpoint** + one-click preset cards (Zhipu, Bailian, Xfyun, ModelScope, SiliconFlow, Qianfan, local Ollama) |
| Transcription cache | Same image + same context reuses the previous transcription — no wasted vision quota (LRU, per-session) |
| Vision fallback chain | Primary vision provider rate-limited/failing → automatic switch to `fallbackProviders` |
| Parallel transcription | `parallelImages` transcribes each image in its own concurrent call (fast multi-image turns) |
| Scene modes | Built-in `transcribeMode` presets: `error-fix` (error-screenshot diagnosis), `chart-sql` (chart → SQL + Pandas), `design-code` (design mockup → HTML+CSS) |
| Cost routing | Small images (≤ `costMaxPixels`) automatically go to a cheap provider |
| Local vision | One-click Ollama preset keeps sensitive images off the network (see [Local vision models](#local-vision-models-keep-sensitive-images-on-your-machine)) |
| Config migration | Settings page exports/imports the whole config as JSON (allow-listed fields only) |

## Security

Since `0.2.1`, the plugin hardens itself against malicious or hand-edited configurations:

- **API key allow-list** (`allowedApiKeyEnvs`): only listed env/credential names may be read as an API key — a malicious config can no longer name `GITHUB_TOKEN` (or any other secret) as `apiKeyEnv` to exfiltrate it
- **Trusted base URLs** (`trustedBaseUrls`): custom vision providers must target official hosts or hosts you explicitly allow — credentials cannot be silently redirected to an attacker endpoint; explicitly listed **local endpoints** (e.g. `http://localhost:11434`) are allowed over plain http and need no key
- **SSRF guard**: generated-image downloads and `reference_image` URLs refuse loopback / private (RFC1918) / link-local addresses
- **No sessionId forwarding**: internal harness session ids are never sent to third-party vision APIs
- **Prompt-injection markers**: vision & OCR outputs are wrapped in explicit "untrusted context" markers before being handed to the text model
- **Sensitive-data redaction** (`redactSensitive`): phone numbers / 18-digit IDs / emails are masked in transcriptions (including cache hits)
- **Audit log** (`auditLog`): one line per transcription with time / image count / bytes / latency / provider
- **Bounded memory**: transcription cache, per-session image recall, and image-kind stats are all LRU-capped — no unbounded growth

## Installation

Prerequisites: official [DeepSeek Harness](https://github.com/deepseek-ai/deepseek-harness) (`0.1.0-rc.6+`) installed and running, Node.js 18+.

```sh
# Option 1: install straight from the GitHub repo (requires git)
dsh plugin --profile web add https://github.com/MC5lan/dsh-multimodal

# Option 2: clone and install from a local path
git clone https://github.com/MC5lan/dsh-multimodal.git
dsh plugin --profile web add /path/to/dsh-multimodal

# Also add for headless mode when needed:
dsh plugin --profile headless add /path/to/dsh-multimodal
```

The repo ships prebuilt artifacts (`lib/`), so you can install right after cloning; to rebuild from source, see [Development](#development).

After restarting `dsh web`, a **「多模态 / Multimodal」** page appears in the Settings sidebar — everything is configured in one place:

- **Status bar** (one row): 识图 ✓/✗ · 画图 ✓/✗ — computed from *real* capability (transcription provider set AND registered / backend exists AND active), so it never shows "configured" when the feature is actually off
- **快速接入** (quick add): the **paste-any-key auto-detect** box is front and center — paste an API key, the plugin figures out the platform (key fingerprint match first, then `/models` probing against known endpoints; key travels one-shot, never stored), and one step later the endpoint is registered, the key env allow-listed, the credential stored, the model list pre-filled, and 识图/画图 enabled. Works with **any OpenAI-compatible API**. Below it: platform preset cards (Zhipu, Bailian, Xfyun, ModelScope, SiliconFlow, Qianfan, local Ollama) and a custom-endpoint form
- **已接入** (connected, open by default): platforms with editable model chips + image-backend management (switch / remove / add form)
- **高级** (advanced, collapsed): the two KeyCards (视觉转写 Key / 图像生成 Key), vision details (watch route, transcription provider, mode, four toggles), and config export/import

Everything in the common path is: paste key → done. The picker then shows
`deepseek-vision` models (e.g. DeepSeek-V4-Flash (Vision)) — that route's
sessions get the "eye".

## Usage

1. **Fastest path (recommended, 0.7+)**: Settings → Multimodal → 快速接入 → paste your API key — platform auto-detected and connected in one step (endpoint + allow-list + credential + model list; 识图/画图 enabled automatically). If the key can't be recognized, the hint says why; then pick a platform card or use a custom endpoint.
2. **Manual path**: tap a preset card (Zhipu / Bailian / Xfyun / ModelScope / SiliconFlow / Qianfan / local Ollama) and paste the key, or declare any OpenAI-compatible endpoint in `extraProviders`; optionally fill `providers.deepseek.models` to advertise models on the built-in DeepSeek route slot.
3. In the Web UI model picker choose a `deepseek-vision` model — that route's sessions get the "eye".
4. Chat normally; paste/drag an image and ask to have it read, or just say "draw me a …" to generate one.

## API Keys (or write the credentials file directly)

The Settings → Multimodal → 密钥管理 page writes both keys for you. To edit
the credentials file directly (path is `$DSH_HOME/.credentials.yaml`, default
`~/.dsh/.credentials.yaml`), it takes effect immediately:

```yaml
VISION_API_KEY: <your-vision-endpoint-key>   # vision transcription (see 视觉转写 Key in the UI)
IMAGE_API_KEY:  <your-image-backend-key>     # image generation (see 图像生成 Key in the UI)
```

> These are *your* provider keys — any OpenAI-compatible endpoint you register
> in 平台接入 accepts its key under the env name you give it (`apiKeyEnv`).

## How it works

```
User image ──▶ host admission (model directory declares image modality → allowed, no popup)
           ──▶ agent/pre-step (plugin): detects images → configured vision engine transcribes
               └─▶ image blocks replaced by "【图片内容转写】…" text
           ──▶ DeepSeek API processes the transcription + your question (image bytes never reach DeepSeek)
```

- The `deepseek-vision` route is the DeepSeek API with a directory that advertises image support so the UI allows attachments — images are actually transcribed by the provider you configured, never sent to DeepSeek
- Image bytes are read through the dsh attachment service; the transcription instruction defaults to "transcribe verbatim + extract key error info" and can be customized in the `dsh-multimodal:` section of `~/.dsh/settings.yaml`
- On transcription failure it degrades to a `【图片转写失败: 原因】` placeholder — DeepSeek keeps going instead of hanging
- Settings UI: Settings → Multimodal (`settings.section` slot, native look & feel like the Models page)

## Configuration (optional, `~/.dsh/settings.yaml`)

```yaml
dsh-multimodal:
  # 0.2.1+ security: only these env/credential names may be used as API keys
  allowedApiKeyEnvs:
    - DEEPSEEK_API_KEY
    # - MY_PROVIDER_API_KEY     # add your custom provider's key env here
  # 0.2.1+ security: extra vision hosts allowed to receive credentials
  trustedBaseUrls:
    # - https://my-vision.example.com   # add your custom provider's host here
  providers:
    deepseek:
      models: []                 # optional: advertise models on the DeepSeek route slot
  vision:
    watchProvider: deepseek-vision  # sessions on this route get the "eye"
    transcribeProvider: ''          # the provider that "sees" (blank = transcription off)
    fallbackProviders: []           # 0.2.2+: providers tried in order when the primary rate-limits/fails
    transcribeMode: auto            # auto | verbatim | structured | ocr | describe | error-fix | chart-sql | design-code
    parallelImages: false           # 0.2.5+: transcribe each image in its own parallel call
    costProvider: ''                # 0.2.5+: cheap provider used for small images (cost routing)
    costModel: ''
    costMaxPixels: 1000000          # images ≤ this many pixels go to costProvider
    sceneHints: true                # 0.2.5+: append diagnosis hint to screenshot transcriptions
    customModes: {}                 # 0.2.4+: custom mode name → prompt pairs (plugin-in transcribers)
    redactSensitive: false          # 0.2.4+: mask phone/ID/email in transcription output
    auditLog: false                 # 0.2.4+: log a line per transcription (time/images/bytes/ms/provider)
  ocr:
    provider: ''                    # OCR tool engine (blank = disabled; any registered vision provider works)
    model: ''
  image:
    backends: {}                    # add backends in the UI; nothing is preloaded
    activeBackend: ''               # image backend key (blank = no image generation)
    failoverOrder: []               # 0.2.5+: backend keys tried in order when the active backend fails
    verifyChineseText: true         # 0.2.3+: check generated Chinese text for garbling via a vision model
    verifyProvider: ''              # vision provider used for the garbling check (blank = check disabled)
  transcribePrompt: ...             # custom transcription instruction
  transcribeTimeoutMs: 90000        # transcription timeout

  # Any OpenAI-compatible vision platform (just add the key; shows up in the model picker & Models page)
  extraProviders:
    xfyun-vision:                   # Xfyun MaaS
      displayName: 讯飞星辰视觉
      baseURL: https://maas-api.cn-huabei-1.xf-yun.com/v2
      apiKeyEnv: XFYUN_API_KEY
      models:
        - id: xoppaddleocrv16       # free OCR (PaddleOCR-VL-1.6)
          name: PaddleOCR-VL-1.6
    siliconflow-ocr:                # SiliconFlow (free hosted DeepSeek-OCR)
      displayName: 硅基流动 OCR
      baseURL: https://api.siliconflow.cn/v1
      apiKeyEnv: SILICONFLOW_API_KEY
      models:
        - id: deepseek-ai/DeepSeek-OCR
          name: DeepSeek-OCR
    modelscope-vision:              # ModelScope (2000 free calls/day)
      displayName: 魔搭视觉
      baseURL: https://api-inference.modelscope.cn/v1
      apiKeyEnv: MODELSCOPE_API_TOKEN
      models:
        - id: Qwen/Qwen3-VL-8B-Instruct
          name: Qwen3-VL-8B
```

Image backends (OpenAI-compatible `/images/generations` and DashScope protocols) example — add them in the UI or write them here:

```yaml
dsh-multimodal:
  image:
    backends:
      z-image:                      # Aliyun Z-Image family (new multimodal-generation sync protocol)
        kind: dashscope-v2
        baseURL: https://dashscope.aliyuncs.com
        apiKeyEnv: DASHSCOPE_API_KEY
        model: z-image-turbo
        defaultSize: 1024*1024      # supports 512*512 ~ 2048*2048
      modelscope-t2i:
        kind: openai-images
        baseURL: https://api-inference.modelscope.cn/v1
        apiKeyEnv: MODELSCOPE_API_TOKEN
        model: <a lightning-badged text-to-image model in the ModelScope plaza>
    activeBackend: z-image
```

### Custom image backends (0.2.6+)

Anything that is not OpenAI-compatible or DashScope can be plugged in with a
small adapter file — no plugin code changes. Write an ES module exporting a
default object with a `generate()` function:

```yaml
dsh-multimodal:
  image:
    backends:
      my-api:
        kind: custom
        adapterFile: D:/my-adapters/my-api.mjs   # absolute or relative path
        baseURL: https://api.example.com/v1       # passed to the adapter
        apiKeyEnv: MY_API_KEY
        model: my-image-model
        defaultSize: 1024*1024
    activeBackend: my-api
    failoverOrder: [z-image]                      # fall back to another configured backend
```

The adapter receives `{ prompt, size, n, negative_prompt, reference_image,
apiKey, baseURL, model, signal, fetch, log }` and returns
`{ urls: string[], b64s: string[] }`. See
[`scripts/adapters/example-custom.mjs`](scripts/adapters/example-custom.mjs)
for a template. **Only point `adapterFile` at files you trust** — the adapter
runs with full process privileges. Polling backends can tune
`pollIntervalMs` / `pollTimeoutMs` per backend.

## Local vision models (keep sensitive images on your machine)

The plugin routes transcription through any OpenAI-compatible endpoint, so a local
[Ollama](https://ollama.com) vision model keeps image data off the network. One-click
preset: Settings → Multimodal → 平台接入 → **+ 本地 Ollama** (auto-adds
`http://localhost:11434/v1` to `trustedBaseUrls`; local endpoints need no API key).

Model matrix (pick by VRAM):

| Model | `ollama pull` | VRAM | Notes |
|---|---|---|---|
| `llava` | `ollama pull llava` | ~4 GB | Classic 7B, good OCR-ish |
| `llava-llama3` | `ollama pull llava-llama3` | ~6 GB | Stronger baseline |
| `qwen2.5vl:7b` | `ollama pull qwen2.5vl:7b` | ~6 GB | Better Chinese text |
| `minicpm-v` | `ollama pull minicpm-v` | ~6 GB | Document/OCR focused |
| `qwen2.5vl:32b` | `ollama pull qwen2.5vl:32b` | ~20 GB | High quality, big VRAM |

Compare providers before wiring them in:

```sh
node scripts/benchmark-vision.mjs shot.png \
  my-endpoint=https://your-vision-api.example.com/v1,your-model,YOUR_API_KEY_ENV \
  ollama=http://localhost:11434/v1,llava,SKIP
```

## Troubleshooting
| Symptom | Cause & fix |
|---|---|
| Reply says "【图片转写失败: MISSING_CREDENTIAL…】" | Vision key missing: add it under Settings → Multimodal → 密钥管理, or write the credentials file |
| Reply says "【图片转写失败: RATE_LIMIT…】" | Free-tier rate limit; retry later or switch to a paid tier |
| Transcription is off (images ignored) | `vision.transcribeProvider` is blank — set it to a registered provider in Settings → Multimodal → 视觉转写 |
| No "Multimodal" entry in Settings | Verify the plugin is mounted (`dsh --profile web --dump-config` should list dsh-multimodal) and hard-refresh the browser (Ctrl+F5) |
| Custom vision provider skipped with a "not trusted" warning | Its base URL isn't allow-listed: add it to `trustedBaseUrls` (or use one of the official hosts); also make sure its key env is in `allowedApiKeyEnvs` |

## Development

```sh
npm install
npm run build      # host: tsc; client: tsc check + tsdown builds lib/client.js
```

## License

[MIT](LICENSE)

Install

dsh plugin --profile web add github:MC5lan/dsh-multimodal#416251d11936ade30c50600eaf5db80ea7b22d21

Profile: web

  • This package builds from source on install. pnpm will ask you to allow its build script — that is permission to run the package’s code on your machine, outside the agent sandbox. Only allow sources you trust.
Source