EdgeBench
eval Your tags
Your notes
Benchmark of 134 day-long real-world agent tasks (51 publicly released) measuring how much agents improve by learning from their environment over 12–72-hour horizons. From ~38,000 hours of runs the team reports that performance follows a log-sigmoid scaling law in interaction time (mean R²=0.998) and that frontier models' "learning speed from environments roughly doubles every three months" — an environment-interaction analogue of the METR long-horizon results. CC-BY-4.0; the report ships as a site PDF, not on arXiv.
Evaluation Details
Tasks 134