Scaling Large Vision-Language Model RL Training via Efficient Load Balancing

Abstract

FlexRL scales reinforcement learning for vision-language models by addressing multimodal data handling and workload imbalance. ShadowLoader distributes visual decoding and preprocessing to workers, while FlexUlysses adaptively shards sequences to balance computation and memory across GPUs.

Publication
In International Conference on Learning Representations (ICLR 2026)
Zerui Wang
Zerui Wang
Ph.D. Student · Research Intern