A coordinated embodied-AI stack release from Ant Group's Robbyant unit (HF org robbyant), July 6–8: four model lines covering the perception→world-model→action pipeline. LingBot-Vision: self-supervised ViT family to ~1B with masked-boundary-modeling for dense spatial perception (Apache-2.0; report pending). LingBot-World 2.0 "Infinity": a 14B causal interactive world model at 720p/60fps with unbounded horizon (CC-BY-NC-SA). LingBot-Video: embodied video-generation FMs (MoE 30B-A3B + 1.3B dense) with a published paper, "Scaling Mixture-of-Experts Video Pretraining for Embodied Intelligence." LingBot-VLA 2.0: a 6B vision-language-action model. Several component reports are still "coming soon" — only the Video paper is on arXiv at filing time.

Outputs 4

LingBot-Vision

model
License Apache 2.0

LingBot-World 2.0

model
Parameters 14B
License CC-BY-NC-SA

LingBot-Video

model
Architecture MOE
Parameters 30B
Active params 3B

LingBot-VLA 2.0

model
Parameters 6B
roboticsembodiedworld-modelopen-weight