"Beyond the Current Observation: Evaluating Multimodal Large Language Models in Controllable Non-Markov Games" — a closed-loop "remember-to-act" evaluation using Reconstructive Non-Markov Games (Matching Pairs and 3D Maze, with a duel protocol), where episodes grow to ~128K tokens and ~350 images. Best frontier results at release: 62.3% on 10x10 Matching Pairs (GPT-5.4) and 50% on the 13x13 maze (Gemini-3.1-Pro). Joint Fudan / Shanghai Innovation Institute / Shanghai AI Laboratory / ZJU / CUHK work (affiliations verified; authors include Haodong Duan, Dahua Lin, Jiaqi Wang). Ships SFT trajectories (rule-based-optimal plus Kimi-K2.5 and Qwen3.5-397B rollouts, MIT).

Paper

evalmultimodalagentsmemory