@yulee-314/dsh-vision-bridge
自包含的 DeepSeek Harness 视觉系统:DeepSeek 视觉孪生路由(原生图片体验 + 视觉桥请求层拦截)+ 本地 Ollama Agentic Vision 工具(describe/OCR/结构化扫描/区域查询/元素定位/双图对比/剪贴板)+ 粘贴分流(paste-to-path)。安装即用,无本机路径依赖。
66 results
自包含的 DeepSeek Harness 视觉系统:DeepSeek 视觉孪生路由(原生图片体验 + 视觉桥请求层拦截)+ 本地 Ollama Agentic Vision 工具(describe/OCR/结构化扫描/区域查询/元素定位/双图对比/剪贴板)+ 粘贴分流(paste-to-path)。安装即用,无本机路径依赖。
Vision for text-only DeepSeek Harness agents: delegate image reading to a one-shot subagent on a configurable vision route (MiniMax / Kimi / any OpenAI-compatible provider), keeping image bytes and the vision model's context out of the main session.
Local-first vision for text-only DeepSeek Harness agents: route image understanding to a local OpenAI-compatible vision model, return structured JSON evidence (summary / OCR / layout / semantics / visual / uncertainty).
Vision 模式:会话区调试/上位机点表(倍率/单位/告警上下限、CSV 导入导出、可视化组件),Keil 编译与日志、产物哈希、OpenOCD 烧录确认、串口报文订阅,Modbus 读点和受控写点(Agent 写点需界面批准),人工操作请求卡、共享任务与时间线,并安装「Vision模式」Agent 预设。
Lightweight, route-preserving vision link for DSH: a configured vision model sees while the selected text model stays in control.
Model-facing analyze_image tool for the DeepSeek Harness: multi-modal image understanding via any OpenAI- or Anthropic-compatible vision API, with 8 analysis modes, local path / http(s) URL / data URL input, and a Web UI hint that guides image-incapable models to the reliable local-path route.
Windows-first vision suite for DeepSeek Harness.
DSH 视觉桥接插件:让无视觉能力的主模型看图(会话收图 + 自动转文字 + view_image 工具)
DSH-native vision bridge for text-only models with native image attachments, multi-image evidence batching, and session-scoped validated Evidence caching.
Gives a text-only LLM vision capability
Adaptive image routing for DeepSeek Harness: pass images directly to native multimodal models and transcribe them only for text-only models.
DSH 静默视觉增强:主模型照常选择,图片自动交给固定视觉模型后以隐藏上下文返回主模型。
DSH 视觉插件(Edge 豆包桥接):通用识图 + 数学建模图专项(几何图形/流程图/图表/表格/公式)+ 不确定项澄清闭环。零成本,免 API Key。
dsh plugin: vision capability proxy. Routes image understanding for text-only main models to a small multimodal model, and backs off entirely when the active model declares multimodal support (inputModalities includes 'image').
DeepSeek Harness plugin: auto-route image-bearing requests to deepseek-v4-flash-vision-exp, then fall back to the original model.
Per-model vision (image input) toggle for DeepSeek Harness (dsh): list every configured model and flip a switch to enable/disable image support without hand-editing settings.yaml. 为 DeepSeek Harness 提供按模型的「支持图片」开关:无需手改 settings.yaml。
DeepSeek Harness plugin: upload/paste images in the chat box; on send, transcribe via a vision model (dashscope) or offline Windows OCR, then inject the description into the message as 【解析了提供图片,图片内容是<描述>】.
DSH vision plugin: image recognition (Codex + Zhipu fallback) and generation (GPT Image + CogView fallback), text-only-model safe
Image generation for DeepSeek Harness: a generate_image tool with pluggable providers — the FAL queue API or any OpenAI-compatible images API. The picture is shown inline in the conversation; the model receives either a link (works with any chat model) or the image itself (needs dsh-vision-bridge or a vision-capable model).
Vision recognition plugin for DeepSeek Harness: paste images into the composer, recognize them via GLM-4V on the host side, and inject the result into the conversation.
DeepSeek Harness plugin: bridge image-bearing messages to a user-configured OpenAI-compatible vision endpoint when the active text model cannot see images
辅助视觉模型:图片→文本描述,供文本模型、浏览器截图兜底与聊天贴图降级使用
Flagship multimodal vision hub for DeepSeek Harness: ~40 tools, PDF drag-and-drop, LaTeX formulas, complex tables, QR codes, UI flow diagrams, and multi-model consensus.
DeepSeek Harness 多模态视觉桥:贴图自动转文字描述(llm/stream 代理)+ view_image/ocr_image 主动视觉工具 + 原生多模态路由自动跳过(rc.7 适配),让 text-only 的 DeepSeek 模型看见图片