Multimodal in-context instruction tuning (legacy anchor, 2023): Otter fine-tuned OpenFlamingo on MIMIC-IT, a 2.8M-sample multimodal instruction dataset with in-context examples, delivering one of the earliest open instruction-tuned LMMs with in-context learning (3.4K★).

Included as the first-in-series artifact of the LMMs-Lab lineage: the Otter/MIMIC-IT team (Bo Li, Yuanhan Zhang, Ziwei Liu at NTU) went on to build LMMs-Eval and the LLaVA-OneVision family. Built on OpenFlamingo, not a from-scratch pretrain.

Paper

multimodalopen-weightresearch

Related