L1 (Length Controlled Policy Optimization)
paper Your tags
Your notes
CMU's LCPO trains reasoning models to hit a target thinking length via RL, giving direct control over the accuracy/compute tradeoff — a widely-adopted lever for efficient reasoning that lets a single model span short and long chains on demand.