Sarashina Image
model Your tags
Your notes
SB Intuitions' first image-generation foundation model, described in a 4 September 2026 research post: a 2.6B-parameter Unified Next-DiT trained from scratch with flow matching, progressive resolution from 256 to 1024, and Flow-GRPO post-training, conditioned by the lab's own Sarashina2.2-Vision-3B text encoder for Japanese-native prompting. Reported results: OneIG 0.497 and UniGenBench 80.42 on long English prompts, ahead of SDXL, FLUX.1-dev, and SD3.5-Large. Weights are not released.
Model Details
Parameters 2.6B