"Benchmarking Long-Horizon Search Agents Beyond the Human Difficulty Ceiling": 544 human-verified questions (plus ~2K synthetic training items) generated from a 7M-entity Wikipedia knowledge graph, positioned as a successor to the saturating BrowseComp. The best frontier model at release (GPT-5.5) scores only 34.7%. MIT license.

Evaluation Details

Questions 544
evalagentssearch

Related