The first real-world benchmark of AI agents exploiting web-application CVEs: critical-severity vulnerabilities reproduced in sandboxed environments with automated success checks. From Daniel Kang's agent-security lab (252★), following its zero-day agent-exploitation line (HPTSA, 2024); now used in agent-security evaluations. The active successor thread to UIUC trustworthy-AI work after Bo Li's DecodingTrust line departed to UChicago / Virtue AI.

Paper

benchmarksafetyagentic