Bridging Agent Semantics with Spot Capacity: An Elastic and Recoverable Service Model
Minchen Yu
Abstract
LLM agents increasingly drive long-running cloud inference workloads in which model calls differ in urgency, redundancy, completion semantics, and replay cost. Model-as-a-Service (MaaS) platforms expose several service models for trading cost against latency, availability, and capacity commitment. These models operate primarily at request, job, or endpoint scopes and provide limited support for combining transient platform supply with the evolving semantics of an agent task. We present SemSpot, a semantics-aware service model that allows agent applications to leverage the spot capacity of LLM inference platforms. At the request level, SemSpot lets a provider publish short-lived offers over successful price, completion probability, and failure-notification deadline; the agent runtime selects among these offers using the current task state and completion rule. An audit of 1,535 cases from six agent benchmarks identifies four recurring workflow structures and shows how this service model may produce different cost, service-time, and fallback behavior. With specialized MaaS support, token-level SemSpot further preserves provider inference state and runtime-verified semantic segments inside a long request. We develop the service model, economic boundary, and the cross-layer research agenda required to realize SemSpot.
Create a lesson
Related papers
Projection-Free Bandit Online Optimization for Multi-Agent Systems with Dynamic Regret
Xia Jiang, Lu Liu, Gang Feng
A Smallest-Need-First Job Scheduling Framework with Adaptive Optimization of Idle Node Counts for Energy-Efficient HPC Systems
Reza Pulungan, Raka Satya Prasasta, Santana Yuda Pradata et al.
CLASP: Chained-Request-Aware Scaling and Operator Placement for Serverless Stream Processing
Tianyu Qi, Maria A. Rodriguez, Rajkumar Buyya
Performance Evaluation of RED-ONION: A High-Speed Disk-to-Disk Transfer System
Keichi Takahashi, Hiroaki Kataoka, Takeo Hosomi et al.
Memory-efficient GPU pipelines for real-time non-line-of-sight reconstruction
Alfonso López-Ruiz, Diego Royo
Great Expectations: Benchmarking the Real-World Performance of RVV 1.0 in HPC
Stepan Nassyr, Prateek Chawla, Daniel Seibel et al.