A diagnostic benchmark (Peking University with Qiyuan Tech) that isolates the effect of the agent harness — the system layer managing context, tools, state, permissions, tracing, and recovery — rather than the base model. Harness-Bench runs representative harness configurations across multiple model backends under shared task environments, budgets, and evaluation protocols, over 106 sandboxed offline tasks vetted for realism, solvability, and oracle-checkability. Across 5,194 execution trajectories it finds substantial variation in completion, process quality, efficiency, and failure mode by model–harness pairing — arguing that agent capability should be reported at the model–harness configuration level, not attributed to the model alone. Surfaces recurring "execution-alignment" failures where plausible reasoning decouples from tool feedback and workspace state.

Paper

evalagentsresearch