Fudan's answer to the compute-matching critique of harness evolution (cf. Rethinking Harness Evolution): propose-and-verify loops waste rollouts scoring every candidate on a fixed task set, and aggregate scores hide specific regressions. HarnessLens makes verification budget-aware and behavior-targeted — it derives candidate modifications from execution trajectories, then selectively verifies each on behavior-relevant tasks behind an attributable-evidence gate.

Reports +7.6–13.6% held-out improvement across three harnesses and four benchmarks at substantially lower evaluation budget. Continues Fudan's agentic-harness-engineering line.

Paper

agentsagent-harnessefficiencyresearch

Related