NeoBabel
model Your tags
Your notes
NeoBabel: A Multilingual Open Tower for Visual Generation — an open multilingual text-to-image foundation model covering six languages, built with large-scale multilingual pretraining plus high-resolution instruction tuning. Reaches 0.75 m-GenEval, state of the art at 2–4x smaller than competing models, with a full open release of code, checkpoints, and a 124M multilingual text-image pair dataset. Led by UvA's Video & Image Sense Lab (Derakhshani first author, Snoek senior); one Cohere Labs coauthor — academic-branded, so it counts as UvA.
Model Details
Architecture DENSE