RLHFlow (RLHF Workflow)
library Your tags
Your notes
"RLHF Workflow: From Reward Modeling to Online RLHF" (TMLR 2024): a fully open, reproducible online-iterative-RLHF recipe — reward and preference model training, iterative DPO, released models and datasets — from Tong Zhang's group with Wei Xiong. The RLHFlow reward models became widely reused defaults in DPO/RLHF research. Edge-of-window preprint (May 2024); published and artifacts released in window.