Tempo-6B
model Your tags
Your notes
"Small VLMs are Smart Compressors for Long Video" — a query-aware long-video multimodal LLM that pairs a Qwen3-VL-2B local compressor with a Qwen3-4B global LLM, letting it reason over hour-plus videos while keeping the token budget tractable. Open weights, three stage checkpoints, and a live demo ship alongside the paper.
The flagship 2026 model of the Mohamed Elhoseiny / Vision-CAIR long-video line, following Goldfish and MiniGPT-4. Because it composes pretrained Qwen3-VL and Qwen3 components rather than pretraining from scratch, its scale is inherited from those bases.