"Harness-Policy Co-Evolution from Agent Experience for Safety Alignment." Shanghai AI Lab's AgentDoG team, with SJTU, Fudan, HKUST, and Zhejiang authors, starts from the observation that an agent's behaviour is shaped jointly by its model and its harness, so safety failures show up in both final responses and multi-step execution. SafeEvolve converts on-policy trajectory safety evidence into bounded, auditable harness updates (a safety prompt plus hierarchical skills), while a two-stage harness-use SFT followed by harness-augmented RL trains the policy to exploit them. Corresponding author Dongrui Liu.

Paper

agentssafetyharness

Related