TULIP: Token-length Upgraded CLIP (ICLR 2025) — breaks CLIP's 77-token caption limit via relative positional encodings plus distillation, with released checkpoints; adopted by the open image-generation community for long-caption retrieval and generation. All-UvA author list (Najdenkoska, Derakhshani, Asano, van Noord, Worring, Snoek); Asano has since left UvA.

Model Details

Architecture DENSE

Paper

Venue ICLR 2025
visionresearch