dsh-eval-harness
DSH 插件回归评测门禁:yaml 用例 + headless 驱动 + trace 断言 + baseline 门禁(eval_run / eval_gate)
12 results
DSH 插件回归评测门禁:yaml 用例 + headless 驱动 + trace 断言 + baseline 门禁(eval_run / eval_gate)
ECC (227k-star operator system) skills for DeepSeek Harness — progressive port of 274 curated single-file skills (agentic engineering, evaluation, testing, patterns, vertical domains, docs). Adapted from affaan-m/ECC (MIT)
AI-driven partial wave analysis for DeepSeek Harness: physics-gated config editing, ctpwa fit execution, numeric evaluation, and goal-driven iterative convergence. Physics knowledge (PDG-2026) as pure functions; DSH integration as a thin pwa_* tool plugin.
Reproducible local experiment matrices for DSH profiles
Benchmark-driven self-evolution plugin for DeepSeek Harness: evaluate, optimize, snapshot, accept or roll back.
Skill-trigger evaluation: an LLM judge recreates the DSH skill catalog and measures how reliably a skill description routes matching queries.
Deterministic, CI-safe golden-output evaluation for DeepSeek Harness
Check which DeepSeek Harness plugins actually loaded, and look up public evaluation boards on trapstreet.run
A local DeepSeek Harness north-star guard with explicit AI indicator evaluation and task alignment context.
Creator-native candidate capture, isolated evaluation, promotion, rollback, and visual lineage for DeepSeek Harness
LLM 评估:输出质量评估、幻觉检测、基准测试、回归守护。受 wshobson/agents(38k★ MIT)启发。
Paired experiments and promotion gates for DSH plugins.