oh-my-knowledge
OMK — Observe. Measure. Know. Evidence-backed knowledge changes for AI applications.
36 results
OMK — Observe. Measure. Know. Evidence-backed knowledge changes for AI applications.
DSH 插件回归评测门禁:yaml 用例 + headless 驱动 + trace 断言 + baseline 门禁(eval_run / eval_gate)
ECC (227k-star operator system) skills for DeepSeek Harness — progressive port of 274 curated single-file skills (agentic engineering, evaluation, testing, patterns, vertical domains, docs). Adapted from affaan-m/ECC (MIT)
Auditable vision and cross-platform Computer Use runtime for DeepSeek Harness with source-preserving evidence.
Tuning Engines CLI, MCP server, and Python agent runtime adapters for governed model, agent, skill, and MCP workflows. Fine-tune open-source LLMs, run inference, manage datasets/evaluations, and connect LangGraph or Temporal while Tuning Engines handles policy, audit, usage, and token economics.
Local-first experiment and evaluation workbench for DeepSeek Harness.
Reproducible local experiment matrices for DSH profiles
Benchmark-driven self-evolution plugin for DeepSeek Harness: evaluate, optimize, snapshot, accept or roll back.
Deterministic, CI-safe golden-output evaluation for DeepSeek Harness
AI-driven partial wave analysis for DeepSeek Harness: physics-gated config editing, ctpwa fit execution, numeric evaluation, and goal-driven iterative convergence. Physics knowledge (PDG-2026) as pure functions; DSH integration as a thin pwa_* tool plugin.
Replay real DeepSeek Harness turns against Standard, Minimal, Anchored, or plugin candidates with frozen request-surface evidence
Skill-trigger evaluation: an LLM judge recreates the DSH skill catalog and measures how reliably a skill description routes matching queries.
Two-stage intent and evidence review, deterministic completion gates, and bounded repair for DeepSeek Harness agents
Audit an agent harness against the harness-evaluation criteria, with machine-enforced evidence validation.
A local DeepSeek Harness north-star guard with explicit AI indicator evaluation and task alignment context.
Compare multiple coding models side by side in DeepSeek Harness with isolated Git worktrees
DSH agent preset for rigorous strategy live-deployment testing/evaluation. Retest.
Community visual workflow and multi-model evaluation plugin for DeepSeek Harness
Check which DeepSeek Harness plugins actually loaded, and look up public evaluation boards on trapstreet.run
Turn explicit coding-agent corrections into executable DeepSeek Harness regression tests.
Replay real DeepSeek Harness turns against Standard, Minimal, Anchored, or plugin candidates with frozen request-surface evidence
Agent benchmark evaluation plugin for DeepSeek Harness: web panel, slash commands, headless CLI, test-set import, model/judge switching, orchestration, scoring, and reports.
Agent observability plugin for DSH — behavior audit, cost tracking, anomaly detection
Evidence packaging, Skill evaluation, and controlled trajectory comparison for DeepSeek Harness sessions