"Bringing Block-wise Diffusion to Vision-Language GUI Agents." A roughly 16.7B MoE vision-language GUI agent whose language backbone is LLaDA2.0-mini-base, so actions are produced by block-wise diffusion decoding rather than autoregression: the GUI observation stays fixed while the action output is progressively denoised. Ant ships the weights, an SGLang server with block-length decoding, and a repository PDF paper with GUI-agent benchmark results; no license is declared on the card.

Model Details

Architecture MOE
Parameters 16.7B
Base model llada-2

Paper

agentsdiffusionmultimodalopen-weight

Related