SEA-LION v4.5
model Your tags
Your notes
The agentic generation of SEA-LION, built by distillation, targeted supervised fine-tuning, and model merging on top of current open models. Qwen-SEA-LION-v4.5-27B-IT post-trains Alibaba's Qwen3.6-27B (hybrid linear and full attention, early-fusion vision, 262K context) on Burmese, Indonesian, Filipino, Malay, Tamil, Thai, and Vietnamese and scores 48 of 107 on the Claw-Eval agentic suite with thinking on, against 47 for the base; Gemma-SEA-LION-v4.5-E2B-IT gives the same treatment to Gemma 4 E2B for on-device use. A SpecDecoder companion, built on z-lab's Qwen3.6-27B-DFlash block-diffusion drafter, is claimed to lift throughput up to five times. All MIT. Not pretrained from scratch.
Model Details
Architecture DENSE
Parameters 27B
Context window 262,144
License MIT
Base model qwen3.6-open
Variants
| Name | Parameters | Notes |
|---|---|---|
| Qwen-SEA-LION-v4.5-27B-IT | 27B | Post-trained from Qwen3.6-27B; Claw-Eval 48/107 with thinking. |
| Gemma-SEA-LION-v4.5-E2B-IT | 2B | Post-trained from Gemma 4 E2B; on-device target. |
| Qwen-SEA-LION-v4.5-27B-IT-SpecDecoder | — | Block-diffusion speculative drafter on Qwen3.6-27B-DFlash; up to 5x throughput claimed. |