The first open-source framework reproducing o1-style reasoning — process reward model training, online/offline RL, and test-time MCTS — from Jun Wang's group (clean UCL). openreasoner/openr reached ~1.8k stars and was widely referenced in the pre-DeepSeek-R1 period as the open reference implementation for reasoning-model training.

Paper

Library

Language Python
infrastructurereasoningreinforcement-learning

Related