Chengxu‘s homepage
I am a training framework engineer on the Dots Infra team at Xiaohongshu, where I work on multimodal training, mentor interns in the RedStar program, and drive research projects within the infrastructure team. My research interests include LLM training infrastructure, pipeline parallelism optimization, and multimodal training.
I received my Ph.D. in Computer Science from Peking University in 2024, advised by Prof. Xuanzhe Liu, and my B.S. in Computer Science from Peking University in 2019.
Email: lh_ycx@126.com
BigMac is a compute- and memory-efficient pipeline system for multimodal LLM training. It nests encoder and generator computation into the LLM pipeline to keep activation memory bounded without sacrificing computational efficiency. Role: Core developer.
UltraEP is a production-ready system for real-time expert load balancing in large-scale MoE training and inference. It dynamically replicates hot experts and reroutes tokens based on the exact load of each microbatch. Role: Contributor.
BigMac: Breaking the Pareto Frontier of Compute and Memory in Multimodal LLM Training. Zili Zhang*, Chengxu Yang*, Shenglong Zhang, Chenyu Wang, Yufan Zhang, Tuo Dai, Zhouyang Li, Yuhong Ge, Chao Jin, Xin Jin, Yuliang Liu. Preprint.
Heddle: A Distributed Orchestration System for Agentic RL Rollout. Zili Zhang, Yinmin Zhong, Chengxu Yang, Chao Jin, Bingyang Wu, Xinming Wei, Yuliang Liu, Xin Jin. Preprint.
UltraEP: Unleash MoE Training and Inference on Rack-Scale Nodes with Near-Optimal Load Balancing. Xinming Wei, Chao Jin, Tuo Dai, Yinmin Zhong, Shan Yu, Chengxu Yang, Bingyang Wu, Zili Zhang, Jing Mai, Qianchao Zhu, Zhouyang Li, Yuliang Liu, Guojie Luo. Preprint.
ReLibra: Routing-Replay-Guided Load Balancing for MoE Training in Reinforcement Learning. Chao Jin, Xinming Wei, Yinmin Zhong, Chengxu Yang, Bingyang Wu, Ruidong Zhu, Zili Zhang, Yuliang Liu, Xin Jin. Preprint.
Adonis: Practical and Efficient Control Flow Recovery through OS-Level Traces. Xuanzhe Liu*, Chengxu Yang*, Ding Li, Yuhan Zhou, Shaofei Li, Jiali Chen, Zhenpeng Chen. ACM Transactions on Software Engineering and Methodology (TOSEM), 2023. [paper] [code]
FLASH: Heterogeneity-Aware Federated Learning at Scale. Chengxu Yang, Mengwei Xu, Qipeng Wang, Zhenpeng Chen, Kang Huang, Yun Ma, Kaigui Bian, Gang Huang, Yunxin Liu, Xin Jin, Xuanzhe Liu. IEEE Transactions on Mobile Computing (TMC), 2022. [paper] [code]
TaintStream: Fine-Grained Taint Tracking for Big Data Platforms through Dynamic Code Translation. Chengxu Yang, Yuanchun Li, Mengwei Xu, Zhenpeng Chen, Yunxin Liu, Gang Huang, Xuanzhe Liu. ESEC/FSE, 2021. [paper] [code]
Characterizing Impacts of Heterogeneity in Federated Learning upon Large-Scale Smartphone Data. Chengxu Yang, Qipeng Wang, Mengwei Xu, Zhenpeng Chen, Kaigui Bian, Yunxin Liu, Xuanzhe Liu. The Web Conference (WWW), 2021. [paper] [code]