BAAI's world foundation model: a unified Next-State-Prediction objective over vision and language (~6B, multimodal encoder-decoder) trained on 125K hours of video and 160M event annotations, doing text generation, next-image prediction, and action generation in one model (text-gen 51.8 avg, image-pred 59.8, robotic-manipulation 32.4 rule-success). Ships with a tech report — BAAI's entry into the world-model race, staged for WAIC.

Model Details

Parameters 6B
world-modelmultimodalopen-weight