ComponentBench: Diagnosing Component-Level Failures in Computer-Use Agents
paper Your tags
Your notes
Amazon AGI SF Lab (senior author Shuyan Zhou) with Duke fills the layer between long-horizon CUA workflow benchmarks and atomic GUI-grounding tests: component-level interactions (e.g. operate a toggle set) that are short enough to diagnose and rich enough to be realistic. ComponentBench instantiates a library-agnostic ontology of 97 canonical UI component types as 2,910 programmatically verified tasks across widely used component libraries, paired with cleaned human reference trajectories so both success and interaction efficiency are measurable, plus a scalable auditing pipeline.