RL in the Wild: Characterizing RLVR Training in LLM Deployment

Abstract

This study characterizes reinforcement learning with verifiable rewards in a deployed LLM training environment. It analyzes workload variation, GPU idling, parallelism, data management, and load imbalance, and introduces the PolyTrace benchmark suite for evaluation with realistic workloads.

Publication
arXiv preprint arXiv:2509.25279
Zerui Wang
Zerui Wang
Ph.D. Student · Research Intern