Representation Finetuning from Stanford NLP (Potts/Wu): intervene on a small set of hidden representations instead of updating weights, matching LoRA at a fraction of the parameters (pyreft 1.5K★). The companion AxBench later showed such steering can beat sparse autoencoders — a load-bearing result for the interpretability-vs- control debate.

Paper

Library

post-trainingresearchinfrastructure