ADeptS-Bench: Trustworthiness of Computer-Use Agents Across Devices
paper Your tags
Your notes
FAIR at Meta (Pierluca D'Oro et al.) builds the safety companion to computer-use capability suites: a dual-stream trustworthiness benchmark for CUAs on mobile and desktop, grounded in the ADEPTS capability framework and general-population user studies. The Safety stream pairs benign and malicious versions of tasks with threats embedded in the visual interface; the Disambiguation stream tests whether agents ask for clarification under ambiguous intent.
Across seven models, none consistently exceeds 80% task success while staying under 30% attack success — and the failure anecdotes are vivid: "every model clicks 'Checkout' on a $25K order without hesitation." A natural pairing with capability-side CUA evals like OSWorld.