Forward-Deployed Full-Stack Engineering for Autonomous Cloud MLOps
cs.MA, cs.AI, cs.LG
Submitted: 2026-08-30
Updated: 2026-08-30
Comments: Nill
License: http://creativecommons.org/licenses/by-nc-sa/4.0/
The gist: Across industries, machine-learning systems support applications ranging from prediction and anomaly detection to forecasting, optimization, and scheduling, yet operationalizing these systems
Terminology
Abstract
Across industries, machine-learning systems support applications ranging from prediction and anomaly detection to forecasting, optimization, and scheduling, yet operationalizing these systems requires coordinating application development, model pipelines, cloud infrastructure, security, deployment, monitoring, retraining, recovery, and rollback. We present an evidence-gated multi-agent framework for transforming a natural-language MLOps cloud engineering task into a verified repository and operational cloud deployment. The framework combines graph engineering, loop engineering, and agent harness engineering. A stateful Graph Orchestrator coordinates specialized agents for repository generation, review, execution, verification, release, and monitoring while governing workflow dependencies, evidence gates, retry bounds, recovery paths, and termination. Consequential lifecycle transitions proceed only when their required predicates are supported by verifiable execution or runtime evidence. Verification failures activate bounded reflection, repair, and re-verification, while runtime evidence of failure, drift, degradation, or policy violation can trigger bounded adaptation, recovery, or rollback. Agent harness engineering constrains repository generation, review, and repair, artifact execution, and cloud operations through controlled capabilities and isolated execution environments. We realize the framework on Google Cloud Platform and evaluate repository completeness, controlled execution, evidence-gated transitions, cloud promotion, and bounded recovery. Our experimental results show that the framework prevents unsupported lifecycle transitions and drives each run toward either a verified operational deployment or an auditable terminal failure.
Sources
- Beyond Task Completion: An Assessment Framework for Evaluating Agentic AI Systems
- Human-on-the-Bridge: Scalable Evaluation for AI Agents
- NL2Repo-Bench: Towards Long-Horizon Repository Generation Evaluation of Coding Agents
- CRITIC: Large Language Models Can Self-Correct with Tool-Interactive Critiquing
- Agentic Software Engineering: Foundational Pillars and a Research Roadmap
- Agentic Harness Engineering: Observability-Driven Automatic Evolution of Coding-Agent Harnesses
- Stop Hand-Holding Your Coding Agent: Engineering the Loops that Replace Step-by-Step Prompting
- What makes prompts a graph: necessary and sufficient conditions for prompt graph engineering
- Towards a Science of AI Agent Reliability
- Reflexion: Language Agents with Verbal Reinforcement Learning
- Time to REFLECT: Can We Trust LLM Judges for Evidence-based Research Agents?
- From Agent Traces to Trust: A Survey of Evidence Tracing and Execution Provenance in LLM Agents
- From Static Templates to Dynamic Runtime Graphs: A Survey of Workflow Optimization for LLM Agents
- AFlow: Automating Agentic Workflow Generation
- Agentic Context Engineering: Evolving Contexts for Self-Improving Language Models
- RepoZero: Can LLMs Generate a Code Repository from Scratch?
- AI Harness Engineering: A Runtime Substrate for Foundation-Model Software Agents
- AgentS4D: Benchmarking Runtime Risks across the Execution Lifecycle of LLM-Based Workspace Agents
- Language Agents as Optimizable Graphs
Related papers
- Highway Congestion Reduction through Reinforcement Learning Based Eulerian Headway Control
- You Only Align Once: Propagating Cooperative Behaviors in Multi-Agent Systems through Seed Agents
- Deny Without Disabling: Authorization-Paired Evaluation and Control for Multi-Agent Systems
- MA-SAPO: Multi-Agent Reasoning for Score-Aware Prompt Optimization
- PeroMAS: A Multi-agent System of Perovskite Material Discovery
- StitchCUDA: An Automated Multi-Agents End-to-End GPU Programing Framework with Rubric-based Agentic Reinforcement Learning