HetRoute Heterogeneous and Cost-aware Collaborative Routing Framework for Distributed Edge MoE Inference
Xin Yuan, Ning Li, Wenchao Xu, Song Guo, Haijun Zhang
Abstract
Mixture-of-Experts (MoE) models have become a dominant architecture for large-scale AI services, yet deploying them over geo-distributed heterogeneous edge servers remains challenging. When the Top-k activated experts of a token are spread across multiple servers, the optimal routing depends jointly on cross-server link bandwidth, heterogeneous GPU computing capability, GPU-CPU expert loading delay, instantaneous queueing backlog, and replica-level quantization quality loss. Existing distributed inference and MoE serving methods address these factors separately and do not provide a unified framework for online multi-server collaborative routing. In this paper, we propose HetRoute, a heterogeneous-cost-aware collaborative routing framework for distributed edge MoE inference. HetRoute introduces a unified per-assignment cost model that explicitly captures four cost components: cross-server transmission, GPU-CPU offloading, GPU computation with queueing, and quantization-induced quality penalty. Guided by this model, the offline stage determines expert server placement, GPU-CPU residency, and replica precision through a routing-cost-coupled deployment algorithm, while the online stage routes the Top-k activated expert set as a whole by minimizing the bottleneck layer cost via exact enumeration or beam search. Theoretical analysis establishes fallback feasibility, a bound on the number of participating servers, per-layer optimality for small candidate domains, and online computational complexity. Trace-driven evaluation on three MoE models over a heterogeneous 10-server edge testbed shows that HetRoute reduces average inference latency by up to 59.0% and P99 latency by up to 58.0%, cuts cross-server traffic by up to 72.1%, and achieves 2.13x throughput improvement compared with representative baselines, while keeping quality degradation within the configured budget.
Create a lesson
Related papers
Taming the Agentic RAN: Stability-Guaranteed Arbitration of Autonomous AI Agents in O-RAN
Seyed Bagher Hashemi Natanzi, Bo Tang
Reliability-Guided Trusted Repeater Node Selection in QKD-Enabled Metro Optical Networks
Arup Kumar Marik, Basabdatta Palit, Sadananda Behera
Toward Composable Network Digital Twins: A Subgraph-Based Latency Prediction Study
Shenjia Ding, David Flynn, Paul Harvey
Jamming Detection in 5G/6G Networks: From O-RAN Concept to OCUDU Deployment
Marcin Hoffmann, Lukasz Kulacz, Osama Baldo et al.
The Operable Pareto Front: Distilling Offline Search into Run-Time Control for Multi-Objective UAV Edge-Computing Scheduling
Qiao Liao, Zhiyong Feng, Bin Wu et al.
Human Exposure to Non-Ionizing Radiation from Indoor Distributed Antenna System: Shopping Mall Measurement Analysis
Júlia da L. A. Silva, Vicente A. de Sousa,, Marcio E. C. Rodrigues et al.