Agent-as-a-Judge proposes using agentic systems to evaluate agentic systems — an evaluator agent that inspects a target agent's intermediate steps, not just its final output — released with the DevAI benchmark of realistic development tasks. The 800★ release has been widely adopted in the agent-evaluation literature.

A joint Meta AI + KAUST AI Initiative effort: KAUST first author Mingchen Zhuge with senior author Jürgen Schmidhuber.

Paper

agentsevaluationresearch