oh-my-knowledge
OMK — Observe. Measure. Know. Evidence-backed knowledge changes for AI applications.
36 results
OMK — Observe. Measure. Know. Evidence-backed knowledge changes for AI applications.
DSH 插件回归评测门禁:yaml 用例 + headless 驱动 + trace 断言 + baseline 门禁(eval_run / eval_gate)
Tuning Engines CLI, MCP server, and Python agent runtime adapters for governed model, agent, skill, and MCP workflows. Fine-tune open-source LLMs, run inference, manage datasets/evaluations, and connect LangGraph or Temporal while Tuning Engines handles policy, audit, usage, and token economics.
ECC (227k-star operator system) skills for DeepSeek Harness — progressive port of 274 curated single-file skills (agentic engineering, evaluation, testing, patterns, vertical domains, docs). Adapted from affaan-m/ECC (MIT)
Local-first experiment and evaluation workbench for DeepSeek Harness.
Best-of-3/5 orchestration and LLM-as-a-Verifier selection for DeepSeek Harness
Reproducible local experiment matrices for DSH profiles
Benchmark-driven self-evolution plugin for DeepSeek Harness: evaluate, optimize, snapshot, accept or roll back.
Deterministic, CI-safe golden-output evaluation for DeepSeek Harness
A blind, fair, local Agent arena inside DSH Web: same task, same commit, isolated worktrees, shared verification, judge before you reveal.
AI-driven partial wave analysis for DeepSeek Harness: physics-gated config editing, ctpwa fit execution, numeric evaluation, and goal-driven iterative convergence. Physics knowledge (PDG-2026) as pure functions; DSH integration as a thin pwa_* tool plugin.
Replay real DeepSeek Harness turns against Standard, Minimal, Anchored, or plugin candidates with frozen request-surface evidence
Skill-trigger evaluation: an LLM judge recreates the DSH skill catalog and measures how reliably a skill description routes matching queries.
Two-stage intent and evidence review, deterministic completion gates, and bounded repair for DeepSeek Harness agents
Audit an agent harness against the harness-evaluation criteria, with machine-enforced evidence validation.
Agent observability plugin for DSH — behavior audit, cost tracking, anomaly detection
A local DeepSeek Harness north-star guard with explicit AI indicator evaluation and task alignment context.
Compare multiple coding models side by side in DeepSeek Harness with isolated Git worktrees
DSH agent preset for rigorous strategy live-deployment testing/evaluation. Retest.
Community visual workflow and multi-model evaluation plugin for DeepSeek Harness
Check which DeepSeek Harness plugins actually loaded, and look up public evaluation boards on trapstreet.run
Turn explicit coding-agent corrections into executable DeepSeek Harness regression tests.
Paired experiments and promotion gates for DSH plugins.
DSH-native multi-runtime baseline, ablation, and reproducible evaluation control plane