LoHoSearch
eval Your tags
Your notes
"Benchmarking Long-Horizon Search Agents Beyond the Human Difficulty Ceiling": 544 human-verified questions (plus ~2K synthetic training items) generated from a 7M-entity Wikipedia knowledge graph, positioned as a successor to the saturating BrowseComp. The best frontier model at release (GPT-5.5) scores only 34.7%. MIT license.
Evaluation Details
Questions 544