Zeppelin: Balancing Variable-length Workloads in Data Parallel Large Model Training

Abstract

Zeppelin balances variable-length workloads in data-parallel large model training. It combines hierarchical sequence partitioning, an attention engine with different parallel strategies, communication routing, and sequence-layout remapping to reduce communication overhead and balance computation.

Publication
In Proceedings of the 21st European Conference on Computer Systems (EuroSys 2026)
Zerui Wang
Zerui Wang
Ph.D. Student · Research Intern