oh-my-knowledge
OMK — Observe. Measure. Know. Evidence-backed knowledge changes for AI applications.
35 results
OMK — Observe. Measure. Know. Evidence-backed knowledge changes for AI applications.
DSH 插件回归评测门禁:yaml 用例 + headless 驱动 + trace 断言 + baseline 门禁(eval_run / eval_gate)
ECC (227k-star operator system) skills for DeepSeek Harness — progressive port of 274 curated single-file skills (agentic engineering, evaluation, testing, patterns, vertical domains, docs). Adapted from affaan-m/ECC (MIT)
Auditable vision and cross-platform Computer Use runtime for DeepSeek Harness with source-preserving evidence.
Tuning Engines CLI, MCP server, and Python agent runtime adapters for governed model, agent, skill, and MCP workflows. Fine-tune open-source LLMs, run inference, manage datasets/evaluations, and connect LangGraph or Temporal while Tuning Engines handles policy, audit, usage, and token economics.
Local-first experiment and evaluation workbench for DeepSeek Harness.
Reproducible local experiment matrices for DSH profiles
Benchmark-driven self-evolution plugin for DeepSeek Harness: evaluate, optimize, snapshot, accept or roll back.
Replay real DeepSeek Harness turns against Standard, Minimal, Anchored, or plugin candidates with frozen request-surface evidence
Two-stage intent and evidence review, deterministic completion gates, and bounded repair for DeepSeek Harness agents
Deterministic, CI-safe golden-output evaluation for DeepSeek Harness
Skill-trigger evaluation: an LLM judge recreates the DSH skill catalog and measures how reliably a skill description routes matching queries.
AI-driven partial wave analysis for DeepSeek Harness: physics-gated config editing, ctpwa fit execution, numeric evaluation, and goal-driven iterative convergence. Physics knowledge (PDG-2026) as pure functions; DSH integration as a thin pwa_* tool plugin.
Audit an agent harness against the harness-evaluation criteria, with machine-enforced evidence validation.
Compare multiple coding models side by side in DeepSeek Harness with isolated Git worktrees
Community visual workflow and multi-model evaluation plugin for DeepSeek Harness
Check which DeepSeek Harness plugins actually loaded, and look up public evaluation boards on trapstreet.run
A local DeepSeek Harness north-star guard with explicit AI indicator evaluation and task alignment context.
DSH-native multi-runtime baseline, ablation, and reproducible evaluation control plane
Replay real DeepSeek Harness turns against Standard, Minimal, Anchored, or plugin candidates with frozen request-surface evidence
A durable, bounded lifecycle supervisor with scheduled evaluation for live DeepSeek Harness sessions
Plugin value auditor for DeepSeek Harness: judge a plugin before install and audit installed ones — heuristic scan + LLM judge, with model-switch re-audit reminders. · DSH 插件价值裁判:装前判断值不值得装,装后审计是否还该留,模型切换时提醒复核。
Agent benchmark evaluation plugin for DeepSeek Harness: web panel, slash commands, headless CLI, test-set import, model/judge switching, orchestration, scoring, and reports.
DeepSeek Harness plugin exposing the CX Agent Studio Scripting API CLI (cxas-scrapi / 'cxas') as model-facing tools: apps, deployments, evaluations, conversations, traces, tools, callbacks, variables, pull/push, and lint.