Scaling Large Vision-Language Model RL Training via Efficient Load Balancing
Zerui Wang, Qinghao Hu, Chang Chen, Jiecheng Zhou, Haojie Duanmu, Xingcheng Zhang, Peng Sun, Dahua Lin
April, 2026
Abstract
FlexRL scales reinforcement learning for vision-language models by addressing multimodal data handling and workload imbalance. ShadowLoader distributes visual decoding and preprocessing to workers, while FlexUlysses adaptively shards sequences to balance computation and memory across GPUs.
Publication
In International Conference on Learning Representations (ICLR 2026)

Ph.D. Student · Research Intern