ML-for-ML
Yutong Zhao, Noga H. Rotman, Gianni Antichi, Ran Ben Basat
Abstract
AI training workloads are growing rapidly, making their time, energy, and infrastructure costs increasingly important. In shared cloud clusters, training and fine-tuning jobs compete with co-running workloads for network resources, while network mechanisms and ML training choices are typically optimized separately: networking controls how bytes move, whereas ML systems control when and how much communication occurs. We argue that this separation leaves end-to-end performance on the table. We present ML-for-ML, a cross-layer perspective in which network-side and ML-side knobs are selected jointly under a shared time-to-target-loss objective. Our preliminary prototype shows that by co-optimizing the ML and network parameters, we reach the target loss up to 42% faster.
Create a lesson
Related papers
Taming the Agentic RAN: Stability-Guaranteed Arbitration of Autonomous AI Agents in O-RAN
Seyed Bagher Hashemi Natanzi, Bo Tang
Reliability-Guided Trusted Repeater Node Selection in QKD-Enabled Metro Optical Networks
Arup Kumar Marik, Basabdatta Palit, Sadananda Behera
Toward Composable Network Digital Twins: A Subgraph-Based Latency Prediction Study
Shenjia Ding, David Flynn, Paul Harvey
Jamming Detection in 5G/6G Networks: From O-RAN Concept to OCUDU Deployment
Marcin Hoffmann, Lukasz Kulacz, Osama Baldo et al.
The Operable Pareto Front: Distilling Offline Search into Run-Time Control for Multi-Objective UAV Edge-Computing Scheduling
Qiao Liao, Zhiyong Feng, Bin Wu et al.
Human Exposure to Non-Ionizing Radiation from Indoor Distributed Antenna System: Shopping Mall Measurement Analysis
Júlia da L. A. Silva, Vicente A. de Sousa,, Marcio E. C. Rodrigues et al.