PairS
paper Your tags
Your notes
Pairwise Preference in LLM Evaluators (COLM 2024): shows LLM-as-judge aligns far better with humans via pairwise comparison plus uncertainty-guided ranking than pointwise scoring — a reference method (~175 citations). Language Technology Lab-led (Zhou, Vulić, Korhonen), with companions ZEPO (EMNLP 2024 Main, a zero-shot prompt optimizer reducing LLM-judge preference bias) and TopViewRS (EMNLP 2024 Oral, exposing VLM weakness at top-view spatial reasoning).