Reverse Engineering Human Preferences with RL
paper Your tags
Your notes
Marek Rei's LAMA lab with Cohere (NeurIPS 2025 Spotlight): RL-trains generators directly against LLM-as-judge preferences, breaking judge-based evaluation pipelines — an adversarial complement to the AISP eval-validity line, and a caution for every leaderboard that relies on model judges.