dsh-llm-multimodal
DSH 插件:提供文本 / 图像 / 视频 / 语音 / 音乐 五个生成工具,模型从 llm-pi-ai 自动发现,含 Settings UI (llm-multimodal namespace)。
153 results
DSH 插件:提供文本 / 图像 / 视频 / 语音 / 音乐 五个生成工具,模型从 llm-pi-ai 自动发现,含 Settings UI (llm-multimodal namespace)。
Automatic reasoning and image capability detection for custom DeepSeek Harness models
DeepSeek Harness 视觉助手插件:给没有视觉能力的模型配一个可随时切换的多模态识别模型。输入框图片自动落盘并改写为文本提示,主模型调用 vision_recognize 工具即可完成看图;识别模型在 settings 的 vision-assist 命名空间热更新切换。
DeepSeek Harness plugin: analyse images out of band with a vision model — pasted images are digested into text before admission and a describe_image tool covers image paths, all without ever changing the session's model.
Give DeepSeek Harness agents eyes: local image analysis via vision models — um_analyze_img tool + model-capability recognition + hot-switch settings UI
Native DSH resource hub: Skills, MCP, third-party plugins and experiments
Declare whether a configured third-party model accepts image (multimodal) input; writes the modality into the owning provider settings and verifies it through runtime model resolution.
Voice for DeepSeek Harness backed by Xiaomi MiMo: browser-native 🎤/🔊 UI (MiMo TTS read-aloud) + voice_transcribe/voice_speak calling MiMo ASR/TTS directly, with a configurable voice map and in-conversation speech strips. Fork of zhuiyueya/dsh-voice (MIT).
Office views for DSH Agent Teams.