"Verifying language model reasoning from the inside out" (ACL 2026 oral; paper November 2025, surfaced with the July 2026 award coverage): a sub-10M-parameter probe over a frozen LLM's internal activations that replaces 1.5–8B-parameter process reward models — 750–810x smaller and 2.6–25x faster in beam-search verification. A strong efficiency datapoint for the process-reward-model line: the reasoning-quality signal is already in the activations.

Paper

Venue ACL 2026 (oral)
reasoningverificationinterpretabilityresearch