Zerui Wang
Zerui Wang
About Me
Publications
Experience
Light
Dark
Automatic
1
Jet-Long: Efficient Long-Context Extension with Dynamic Bifocal RoPE
Jet-Long is a tuning-free method for extending the context length of large language models. It combines a local RoPE-faithful window with a dynamically rescaled long-range window and an efficient attention implementation, preserving short-context behavior while supporting longer inputs without retraining.
Haozhan Tang
,
Zerui Wang
,
Yuxian Gu
,
Song Han
,
Han Cai
PDF
Cite
Code
DOI
arXiv
Zeppelin: Balancing Variable-length Workloads in Data Parallel Large Model Training
Zeppelin balances variable-length workloads in data-parallel large model training. It combines hierarchical sequence partitioning, an attention engine with different parallel strategies, communication routing, and sequence-layout remapping to reduce communication overhead and balance computation.
Chang Chen
,
Tiancheng Chen
,
Jiangfei Duan
,
Qianchao Zhu
,
Zerui Wang
,
Qinghao Hu
,
Peng Sun
,
Xiuhong Li
,
Chao Yang
,
Torsten Hoefler
PDF
Cite
DOI
arXiv
Scaling Large Vision-Language Model RL Training via Efficient Load Balancing
FlexRL scales reinforcement learning for vision-language models by addressing multimodal data handling and workload imbalance. ShadowLoader distributes visual decoding and preprocessing to workers, while FlexUlysses adaptively shards sequences to balance computation and memory across GPUs.
Zerui Wang
,
Qinghao Hu
,
Chang Chen
,
Jiecheng Zhou
,
Haojie Duanmu
,
Xingcheng Zhang
,
Peng Sun
,
Dahua Lin
PDF
Cite
Proceedings
PackMamba: Efficient Processing of Variable-Length Sequences in Mamba Training
PackMamba improves the efficiency of training Mamba models on variable-length sequences. It packs sequences and modifies the convolution and selective-scan operators to preserve sequence boundaries, reducing padding overhead while maintaining parallel execution.
Haoran Xu
,
Ziqian Liu
,
Rong Fu
,
Zhongling Su
,
Zerui Wang
,
Zheng Cai
,
Zhilin Pei
,
Xingcheng Zhang
PDF
Cite
Code
DOI
arXiv
SchedMate: Large Language Models are DL Scheduling Enhancers
Existing deep learning cluster schedulers make their scheduling decisions on partial job information, such as resource utilization …
Zerui Wang
,
Qinghao Hu
,
Ana Klimovic
,
Xingcheng Zhang
,
Peng Sun
PDF
Cite
Characterization of Large Language Model Development in the Datacenter
Large Language Models (LLMs) have presented impressive performance across several transformative tasks. However, it is non-trivial to …
Qinghao Hu
,
Zhisheng Ye
,
Zerui Wang
,
Guoteng Wang
,
Meng Zhang
,
Qiaoling Chen
,
Peng Sun
,
Dahua Lin
,
Xiaolin Wang
,
Yingwei Luo
PDF
Cite
Code
Dataset
Cite
×