Small fact-checkers for grounding LLM output on documents (EMNLP 2024): GPT-4-level verification accuracy at ~1/400 of the cost, trained on synthetic error data. Adopted well beyond the paper — Ollama ships Bespoke-MiniCheck, and Guardrails AI integrates it as a grounding validator. From Greg Durrett's group; attribution stays with UT for this work, though Durrett's TAUR lab has since moved to NYU.

Paper

Library

evaluationefficiencyopen-source