"Evaluating LLM Agents on Long-Horizon Autonomous Business Operations." A Qwen Team benchmark, built with Taobao and Tmall and HKUST, in which an agent is handed ¥100,000 and up to four stores and must run them for 365 simulated days inside a deterministic market with multi-round supplier negotiation and dynamic events such as promotions, disasters, and supply shocks; the primary metric is the asset multiplier over five episodes. Eighteen models sit on the launch leaderboard. Posted August 31, 2026 with an Apache 2.0 repository (85 stars). Filed late.

Paper

Evaluation Details

Scoring asset multiplier after 365 simulated days, averaged over 5 episodes
evalagentslong-horizon

Related