This study characterizes reinforcement learning with verifiable rewards in a deployed LLM training environment. It analyzes workload variation, GPU idling, parallelism, data management, and load imbalance, and introduces the PolyTrace benchmark suite for evaluation with realistic workloads.