A native multimodal full-duplex interaction model from Tencent's Hunyuan Speech Team (with NTU), built on OpenBMB's MiniCPM-o 4.5 and adapted for real-time use with an asynchronous agent loop. Instead of turn-taking, Gander streams video, speech, and text continuously: the user can interrupt, and the model can volunteer intermediate feedback or ask follow-up questions mid-task. Its Cerebellum–Brain design splits a real-time interaction module from a slower reasoning module, so complex workflow agents run behind a responsive conversational front. Released 8 September 2026 as a 9B model with code on GitHub.

Model Details

Parameters 9B
Base model minicpm-o4.5

Paper

multimodalspeechagenticopen-weight

Related