Co-Evolving Harnesses and Models
paper Your tags
Your notes
"On-Policy Correction Helps Weaker Models Catch Up Where Imitation Fails." Salesforce AI (Zhou Yu, Bin Bi, Ran Xu and colleagues) evolves a harness with a weak model across seven enterprise agent tasks, shows that a stronger expert uses the evolved harness better still, and that imitating the expert breaks the harness fit, whereas on-policy correction closes the gap. A companion to EvoHarnessBench and to KRAFTON's WHALE on joint harness-weight optimization.