Natural Language Reinforcement Learning
paper Your tags
Your notes
NLRL redefines the RL primitives — value, the Bellman equation, policy iteration — in natural-language space, a conceptual foundation cited across LLM-agent RL work. Clean UCL (Feng/Wang, Jun Wang group); companion to the group's OpenR reasoning infrastructure.