A unified benchmark and framework for LLM unlearning from Kolter's locuslab (NeurIPS 2025 Datasets & Benchmarks; 568★), consolidating the TOFU/MUSE lineage into a standard evaluation substrate — the reference toolkit for measuring whether models have actually forgotten targeted data.

Paper

Library

safetyevalinfrastructure