Optimizing Effective Training Time for Large-Scale Recommendation Systems
Mingming Ding, Ruilin Chen, Yuzhen Huang, Hang Qi, Menglu Yu, San Tan, Damian Reeves, Boris Sarana, Kevin Tang, Satendra Gera, Gagan Jain, Sahil Shah, Vishwa Karia, Fuzail Khan, Yashasvi Makin, Edward Z. Yang, Oguz Ulgen, Jia Chen Ren, Laith Sakka, Mayank Garg, Meet Vadakkanchery, Aici Lin, Wei Sun, Mengjiao Zhou, Shuai Yang, Junqing Zhou, Max Leung, Apoorv Purwar, Musharaf Sultan, John Bocharov, Zhenyu Tang, Vivek Trehan
Abstract
Lifecycle overhead silently consumes accelerator capacity across large-scale recommendation training fleets. Our largest recommendation workloads process tens of billions train- ing examples per day on thousands of GPUs. Before this work, only 50-60% of their end-to-end wall time advanced training on new data. We present a fleet-scale study of this lifecycle overhead and a set of optimizations spanning the full training stack. We use Effective Training Time (ETT%) as an operational framework to instrument lost time, localize it to independently owned infrastructure components, and expose work repeated across job restarts. This analysis guides optimizations like communication elimination and pipeline overlap during trainer initialization; dynamic-shape handling, autotuning pruning, and reusable Py- Torch 2 compilation caches; asynchronous checkpointing; stan- dalone model publishing; and reductions in recovery cost. We evaluate the optimizations on representative models and measure their impacts in our training fleet. ETT% improves on every benchmark, by 15.5% on average, and reaches 85% on our largest workload. Fleet-wide ETT% rose from about 80% to above 90% after deployment.
Create a lesson
Related papers
AgentWebRec: Compact Evidence Fusion over the Agent Web for Personalized Recommendation
Haoran Qiang, Guannan Liu, Liang Zhang et al.
From Rules to Neural Graphs: Scalable Structured Prediction for Patent Prior Art Search
Nikolai Zenovkin, Sebastian Björkqvist
Learning to structure data from user-generated thematic corpora
Elishay Avram, Oren Glickman, Elad Yom-Tov
Do Multilingual Encoders Produce Language-Consistent Semantic IDs?
Abhinav Bohra, Anuj Bohra
RPTune: Learned Context Curation for LLM Catalog Search
Chuxuan Hu, Hejie Cui, Norman Huang et al.
CANOPY: Adaptive-Granularity Evidence Compression for Multimodal RAG
Hyojeong Yun, Jueun Kim, Wook-Shin Han