oh-my-knowledge
OMK — Observe. Measure. Know. Evidence-backed knowledge changes for AI applications.
36 results
OMK — Observe. Measure. Know. Evidence-backed knowledge changes for AI applications.
DSH 插件回归评测门禁:yaml 用例 + headless 驱动 + trace 断言 + baseline 门禁(eval_run / eval_gate)
ECC (227k-star operator system) skills for DeepSeek Harness — progressive port of 274 curated single-file skills (agentic engineering, evaluation, testing, patterns, vertical domains, docs). Adapted from affaan-m/ECC (MIT)
Auditable vision and cross-platform Computer Use runtime for DeepSeek Harness with source-preserving evidence.
Tuning Engines CLI, MCP server, and Python agent runtime adapters for governed model, agent, skill, and MCP workflows. Fine-tune open-source LLMs, run inference, manage datasets/evaluations, and connect LangGraph or Temporal while Tuning Engines handles policy, audit, usage, and token economics.
Local-first experiment and evaluation workbench for DeepSeek Harness.
Reproducible local experiment matrices for DSH profiles
Benchmark-driven self-evolution plugin for DeepSeek Harness: evaluate, optimize, snapshot, accept or roll back.
AI-driven partial wave analysis for DeepSeek Harness: physics-gated config editing, ctpwa fit execution, numeric evaluation, and goal-driven iterative convergence. Physics knowledge (PDG-2026) as pure functions; DSH integration as a thin pwa_* tool plugin.
Replay real DeepSeek Harness turns against Standard, Minimal, Anchored, or plugin candidates with frozen request-surface evidence
Skill-trigger evaluation: an LLM judge recreates the DSH skill catalog and measures how reliably a skill description routes matching queries.
Audit an agent harness against the harness-evaluation criteria, with machine-enforced evidence validation.
Compare multiple coding models side by side in DeepSeek Harness with isolated Git worktrees
DSH agent preset for rigorous strategy live-deployment testing/evaluation. Retest.
Deterministic, CI-safe golden-output evaluation for DeepSeek Harness
Community visual workflow and multi-model evaluation plugin for DeepSeek Harness
Check which DeepSeek Harness plugins actually loaded, and look up public evaluation boards on trapstreet.run
Two-stage intent and evidence review, deterministic completion gates, and bounded repair for DeepSeek Harness agents
A local DeepSeek Harness north-star guard with explicit AI indicator evaluation and task alignment context.
Controlled A/B comparisons and evidence-backed reports for DeepSeek Harness plugins and presets.
Turn explicit coding-agent corrections into executable DeepSeek Harness regression tests.
Replay real DeepSeek Harness turns against Standard, Minimal, Anchored, or plugin candidates with frozen request-surface evidence
Agent benchmark evaluation plugin for DeepSeek Harness: web panel, slash commands, headless CLI, test-set import, model/judge switching, orchestration, scoring, and reports.
Agent observability plugin for DSH — behavior audit, cost tracking, anomaly detection