AceSpec: An Asymmetric Edge-Cloud Collaborative Framework for Communication-Efficient LLM Inference
Yida Zhang, Zhiyong Gao, Shuaibing Yue, Jie Li, Rui Wang
Abstract
Deploying Large Language Models (LLMs) on edge devices typically relies on model compression or split inference. However, compression degrades reasoning capabilities, while split inference suffers from severe Wide Area Network (WAN) communication bottlenecks. Edge-cloud speculative decoding emerges as a promising alternative, leveraging an edge small model to draft tokens for cloud verification. Yet, over volatile WANs, inevitable prediction rejections trigger catastrophic pipeline stalls and network-wide rollbacks, neutralizing collaborative gains. To overcome this, we propose AceSpec, an asymmetric edge-cloud collaborative framework. AceSpec utilizes un-saturated edge compute to proactively construct a probabilistic state cache, effectively transforming network-wide pipeline flushes into O(1) local memory lookups. To preserve bandwidth, it employs an asymmetric communication protocol that transmits minimal main-chain indices uplink and compact sparse distributions downlink. Furthermore, we introduce a network-aware, Lagrangian-optimized resource allocation strategy that dynamically maximizes the local cache hit rate. Evaluations demonstrate that AceSpec achieves up to a 3.52× throughput speedup and exhibits exceptional bandwidth immunity, sustaining near-peak inference performance even under severely constrained 50 Kbps WAN conditions.
Create a lesson
Related papers
Federated Learning on the American Science Cloud using APPFL
Zilinghan Li, Abhijit Chunduru, Harinarayan Krishnan et al.
Towards Global Federated Genome-Wide Association Meta-Analysis Using GA4GH TES
Abhijit Chunduru, Matthew Joel, Zilinghan Li et al.
MeanField Surrogate Modeling for Scalable Runtime Scheduling of Concurrent Heterogeneous AI Inference on Shared GPUs
Youssef Ennouri, Soonhoi Ha
RT-HiSS: Ray Tracing Accelerated High Dimensional Vector Similarity Searches
Revanth Reddy Munugala, Michael Gowanlock
CREDIT: Cost-guided Reduction-reuse with Efficient DSMEM Inter-CTA Tiling
Zhengxiong Li, Tsung-Wei Huang, Umit Ogras
Scaling Inference Prefill with High-Radix Photonic Interconnects
Arulselvan Madhavan, Peter Carson, Taylor Groves et al.