dsh-eval-harness
DSH 插件回归评测门禁:yaml 用例 + headless 驱动 + trace 断言 + baseline 门禁(eval_run / eval_gate)
15 results
DSH 插件回归评测门禁:yaml 用例 + headless 驱动 + trace 断言 + baseline 门禁(eval_run / eval_gate)
ECC (227k-star operator system) skills for DeepSeek Harness — progressive port of 274 curated single-file skills (agentic engineering, evaluation, testing, patterns, vertical domains, docs). Adapted from affaan-m/ECC (MIT)
AI-driven partial wave analysis for DeepSeek Harness: physics-gated config editing, ctpwa fit execution, numeric evaluation, and goal-driven iterative convergence. Physics knowledge (PDG-2026) as pure functions; DSH integration as a thin pwa_* tool plugin.
Reproducible local experiment matrices for DSH profiles
Benchmark-driven self-evolution plugin for DeepSeek Harness: evaluate, optimize, snapshot, accept or roll back.
Skill-trigger evaluation: an LLM judge recreates the DSH skill catalog and measures how reliably a skill description routes matching queries.
Deterministic, CI-safe golden-output evaluation for DeepSeek Harness
Check which DeepSeek Harness plugins actually loaded, and look up public evaluation boards on trapstreet.run
Installable bundle contributing the DaoZang offline retrieval & original-text extraction skill to DeepSeek Harness (v2.0: launcher + self-check + workspace setup)
Local-first, source-linked Markdown wiki storage and retrieval plugin for DeepSeek Harness
A local DeepSeek Harness north-star guard with explicit AI indicator evaluation and task alignment context.
把技能当可训练参数:像训练神经网络一样训练 agent 技能(epochs/batchsize/learning rate/验证门禁,但不碰模型权重)——rollout→reflect→aggregate→select→update→evaluate 循环、候选编辑仅在严格改善 held-out 验证分时接受、文本学习率预算、零推理时模型调用、部署紧凑 best_skill.md。受 microsoft/SkillOpt(MIT)启发。
Creator-native candidate capture, isolated evaluation, promotion, rollback, and visual lineage for DeepSeek Harness
LLM 评估:输出质量评估、幻觉检测、基准测试、回归守护。受 wshobson/agents(38k★ MIT)启发。
Paired experiments and promotion gates for DSH plugins.