dsh-design-qa
让 DeepSeek Harness 里任何纯文本模型都能读图。识图是按需调用的 tool —— 图片不进主模型上下文,不看就不花钱;附 23 处缺陷的评测集,换模型可自测。
24 results
让 DeepSeek Harness 里任何纯文本模型都能读图。识图是按需调用的 tool —— 图片不进主模型上下文,不看就不花钱;附 23 处缺陷的评测集,换模型可自测。
OMK — Observe. Measure. Know. Evidence-backed knowledge changes for AI applications.
Closed-loop trajectory policy plane for DeepSeek Harness: task episodes, revision-aware verification debt, benchmark-aware stop control, adaptive effort, and scoped capability control.
Local-first experiment and evaluation workbench for DeepSeek Harness.
Benchmark-driven self-evolution plugin for DeepSeek Harness: evaluate, optimize, snapshot, accept or roll back.
DeepSeek Harness tools for reproducing the ml-quant-trading protocol v1 benchmark.
Native continuity recovery for long-running DeepSeek Harness agents
Reproducible deterministic benchmark evidence for DSH tools and plugins
Reproducible local experiment matrices for DSH profiles
A reusable DeepSeek Harness bundle for evidence-driven memory, orchestration, benchmark operations, repository audits, and plugin release workflows.
A blind, fair, local Agent arena inside DSH Web: same task, same commit, isolated worktrees, shared verification, judge before you reveal.
经验库(更有经验的DeepSeek):实时采集/打标队列/半自动+全自动空闲派活/位置索引/10本试验经验技能书+benchmark数据。适配 meow-memory 教训层。host 半 esbuild 自包含产物 + export default 插件形态。
Model capability radar for the DeepSeek Harness web GUI: a Settings tab that charts codexradar crowd-benchmark scores per model tier — task-composition bars, 7-day IQ trend, efficiency badges
Run one command N rounds and judge by median/distribution instead of a single run
Check which DeepSeek Harness plugins actually loaded, and look up public evaluation boards on trapstreet.run
灰度模型检测工具:一键开测 40 个并发 dsh 会话(V4 Pro + max 思考),流式监听思维链——出现 I'm/I'll 特征标记为灰度神模型,出现 Let me 标记为正常模型;统计灰度出现率,结果可在设置页查看。
A deterministic piano performance plugin for DSH: shared musical timeline, piano audio engine, visual engine, and DSH-compatible host/browser boundary.
DSH plugin that bundles the dsh-skill-creator skill: create, benchmark, review, and iterate DSH skills.
Paired experiments and promotion gates for DSH plugins.
Agent benchmark evaluation plugin for DeepSeek Harness: web panel, slash commands, headless CLI, test-set import, model/judge switching, orchestration, scoring, and reports.
Talk to Excel in DeepSeek Harness: create/edit spreadsheets, formulas, styles, filters, and tables by conversation; auto-verifies formulas after every edit, repairs silent errors, and validates charts.
Performance diagnosis, repeated plugin isolation campaigns, safe recovery, and measured verification for DeepSeek Harness
Controlled A/B comparisons and evidence-backed reports for DeepSeek Harness plugins and presets.
Per-vendor/per-model LLM latency telemetry and cross-vendor benchmark plugin for DeepSeek Harness