LLMs Know More Than They Show
paper Your tags
Your notes
"LLMs Know More Than They Show: On the Intrinsic Representation of LLM Hallucinations" (ICLR 2025): the truthfulness signal inside an LLM is concentrated in specific tokens, and exploiting this enables sharply better error detection than reading the model's surface outputs. Widely adopted (the LLMsKnow repo, ~95★). By Yonatan Belinkov's group, with some co-authors at Google/Apple.