NVIDIA systems work on the rollout-cost axis of large-scale, long-context RL post-training: a speculative-decoding draft model is co-trained online with the policy, with branch attention under zigzag ring context parallelism and cross-pipeline feature transport so the draft keeps pace as the policy changes.

Paper

rlinfrastructureinference