Online Draft Co-Training for Speculative Decoding in RL Post-Training
paper Your tags
Your notes
NVIDIA systems work on the rollout-cost axis of large-scale, long-context RL post-training: a speculative-decoding draft model is co-trained online with the policy, with branch attention under zigzag ring context parallelism and cross-pipeline feature transport so the draft keeps pace as the policy changes.