dsh-design-qa
让 DeepSeek Harness 里任何纯文本模型都能读图。识图是按需调用的 tool —— 图片不进主模型上下文,不看就不花钱;附 23 处缺陷的评测集,换模型可自测。
12 results
让 DeepSeek Harness 里任何纯文本模型都能读图。识图是按需调用的 tool —— 图片不进主模型上下文,不看就不花钱;附 23 处缺陷的评测集,换模型可自测。
OMK — Observe. Measure. Know. Evidence-backed knowledge changes for AI applications.
A blind, fair, local Agent arena inside DSH Web: same task, same commit, isolated worktrees, shared verification, judge before you reveal.
Replay real DeepSeek Harness turns against Standard, Minimal, Anchored, or plugin candidates with frozen request-surface evidence
Two-stage intent and evidence review, deterministic completion gates, and bounded repair for DeepSeek Harness agents
Compare multiple coding models side by side in DeepSeek Harness with isolated Git worktrees
Community visual workflow and multi-model evaluation plugin for DeepSeek Harness
Paired experiments and promotion gates for DSH plugins.
Replay real DeepSeek Harness turns against Standard, Minimal, Anchored, or plugin candidates with frozen request-surface evidence
A proof-carrying correction loop for DSH: explicit adoption, scoped recall, and reconstructable delivery.
Controlled A/B comparisons and evidence-backed reports for DeepSeek Harness plugins and presets.
LLM 评估:输出质量评估、幻觉检测、基准测试、回归守护。受 wshobson/agents(38k★ MIT)启发。