MoSE: Mode-Switching Expander for Mixed LLM Training and Inference
Fan Yang, Ying Zhou, Binglei Wang, Zhenjie Zhou, Jialong Li
Abstract
AI clusters increasingly run large language model (LLM) inference and training on the same fabric. Prefill-decode (P-D) disaggregation creates key-value (KV) cache transfers between prefill and decode groups, whereas training collectives and all-to-all traffic benefit from near-uniform global connectivity. A static sparse topology can therefore be poorly matched to one of the two traffic patterns. We present Mode-Switching Expander (MoSE), a reconfigurable expander that treats topology design as a fixed-degree edge-allocation problem. MoSE reallocates the same sparse edge budget toward direct P-D connectivity in inference-heavy modes and restores a uniform random regular expander in training-heavy modes. We evaluate MoSE using a 1024-group flow-level topology model, shortest-path routing, and two mixed workloads. Across 20 seeds, MoSE reduces average and 95th-percentile (P95) load-aware KV communication cost by 90.8\% and 91.9\% relative to Static-Training in the inference-heavy mode. In the training-heavy mode, it reduces average and P95 training communication cost by 22.7\% and 27.6\% relative to stale Static-Inference. These results show that coarse-grained topology switching can support both workload modes without additional ports or routing changes.
Create a lesson
Related papers
SL-RFSIM: Enabling Scalable Multi-Hop 5G NR Sidelink Mesh Networking in OpenAirInterface
Simone Pio Candido, Jin Yan, Jérôme Härri
DRL-driven RAN Slicing Management: A V2X-oriented Approach In Multi-service Scenarios
Daniel E. Garcia-Fernandez, Pablo Vera-Soto, Sergio Fortes et al.
Federated Learning for LLMs over Mobile Networks: Issues and Solutions in the RAN Transport
Emilio Paolini, Andrea Pinto, Flavio Esposito et al.
Resource-Efficient Semantic Communication for Heterogeneous Agentic Teams
Farhad Rezazadeh, Hatim Chergui, Lingjia Liu et al.
Spatio-Temporal Wireless-Optical Planning for Multi-UAV Networks
Binglei Wang, Huiru Ao, Fan Yang et al.
Can LLMs help find Ambiguities in Protocol Specifications?
Ziyue Dang, Sixu Tan, Atharva Nevasekar et al.