Forward-Deployed Full-Stack Engineering for Autonomous Cloud MLOps
Sagar Srinivas Sakhinana, Venkataramana Runkana
Abstract
Across industries, machine-learning systems support applications ranging from prediction and anomaly detection to forecasting, optimization, and scheduling, yet operationalizing these systems requires coordinating application development, model pipelines, cloud infrastructure, security, deployment, monitoring, retraining, recovery, and rollback. We present an evidence-gated multi-agent framework for transforming a natural-language MLOps cloud engineering task into a verified repository and operational cloud deployment. The framework combines graph engineering, loop engineering, and agent harness engineering. A stateful Graph Orchestrator coordinates specialized agents for repository generation, review, execution, verification, release, and monitoring while governing workflow dependencies, evidence gates, retry bounds, recovery paths, and termination. Consequential lifecycle transitions proceed only when their required predicates are supported by verifiable execution or runtime evidence. Verification failures activate bounded reflection, repair, and re-verification, while runtime evidence of failure, drift, degradation, or policy violation can trigger bounded adaptation, recovery, or rollback. Agent harness engineering constrains repository generation, review, and repair, artifact execution, and cloud operations through controlled capabilities and isolated execution environments. We realize the framework on Google Cloud Platform and evaluate repository completeness, controlled execution, evidence-gated transitions, cloud promotion, and bounded recovery. Our experimental results show that the framework prevents unsupported lifecycle transitions and drives each run toward either a verified operational deployment or an auditable terminal failure.
Create a lesson
Related papers
Harness-RL: Black-Box Reinforcement Learning with Action-Args Decoupling for Central-Agent Multi-Agent Harnesses
Xinke Jiang, Zhixin Zhang, Zhibang Yang et al.
AgenticRag-R1: Agentic Reinforcement Learning with Stack Memory for Multi-Step Reasoning, Retrieval and Memorizing
Xinke Jiang, Yue Fang, Zhibang Yang et al.
Generalizable Multi-Agent Planning from Signal Temporal Logic Specifications via Diffusion
Joe Eappen, Zikang Xiong, Shreyash S. Iyengar et al.
ASTRA - Agentic System for Ticket Resolution and Analysis
Shashidhar Reddy Javaji, Mohamed Trabelsi, Jin Cao et al.
One Model, Many Minds: Unlocking Multi-Agent Synergy in a Single Agent via Mixture of Roles
Zhichen Zeng, Huiyuan Chen, Jingru Cheng et al.
Dynamic Haven Selection for Multi-Agent Pickup and Delivery in Constrained Warehouses
Taisei Hirayama, Kohei Yoshida, Hiroki Sakaji et al.