The vision-language member of the LLM-jp-4 generation: NII's own 8.6B llm-jp-4-8b-thinking backbone joined to a SigLIP2 so400m encoder (0.4B) through a two-layer MLP projector, about 9B in all. Trained on roughly 29.3M samples, 9.2M Japanese (Jagle), 12.0M English (RefinedVision, released with the model), a 5.0M subset of Nemotron-Image-Training-v3, and 3.2M thinking-SFT examples; a 25-benchmark average of 61.0 against 58.3 for the earlier beta. Apache 2.0. The language backbone is NII's from-scratch model; only the vision encoder is borrowed.

Model Details

Architecture DENSE
License Apache-2.0
Base model llm-jp-4
multimodaljapaneseopen-weight

Related