Benchmarking Agentic LLM/VLM Reasoning On Games (ICLR 2025): rolls NetHack, MiniHack, Crafter, and TextWorld into the standard long-horizon LLM-agent benchmark, with external adoption including NVIDIA NIM benchmarking. UCL DARK (Paglieri, Rocktäschel).

Paper

Venue ICLR 2025

Evaluation Details

Domains 3
Scoring agentic task success
Domains: games, agents, reasoning
evalagentsreasoning