SEA-HELM
evalYour notes
The regional benchmark the SEA-LION models are built against, developed by AI Singapore with Stanford CRFM's HELM team and published at ACL 2025 Findings (Yosephine Susanto first author, William Chandra Tjhi senior author). SEA-HELM scores models along five pillars, NLP classics, LLM-specific tasks such as instruction following and chat, Southeast Asian linguistics (the LINDSEA diagnostics), SEA culture, and safety, in Filipino, Indonesian, Tamil, Thai, and Vietnamese at launch, with about 35,000 items across translation, reasoning, understanding, toxicity, and cultural tests; the 2026 revision used for v4.8 adds Malay and Burmese and folds in SEA-SafeguardBench. Its public leaderboard ranks open and closed models by language and task and is the source of the SEA-HELM numbers quoted across the SEA-LION entries.