Towards Agentic Cloud Engineering: Graph and Loop Engineering with a Zero-Trust Agent Harness
Listen
Radio episode about this paper
Transcript
Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.
Tom: Next we'll be talking about the paper "Towards Agentic Cloud Engineering: Graph and Loop Engineering with a Zero-Trust Agent Harness".
Jane: The paper was written by Sagar Srinivas Sakhinana and Venkataramana Runkana from Tata Research Development and Design Centre.
Tom: Stay tuned as we take you through the paper and discuss its implications.
Summary: Tom: We’ve established that the title suggests a highly structured, trustworthy approach to AI-driven cloud tasks, but let's look at the summary of "Towards Agentic Cloud Engineering: Graph and Loop Engineering with a Zero-Trust Agent Harness."
Jane: The authors summarize this by showing how it transforms a natural language request into three things: a validated code repository, a verified operational cloud deployment, and the verification evidence.
Lu: This moves beyond mere task completion; we are talking about creating artifacts that are verifiable by machine logic itself, which is incredibly powerful.
Meng: I want to understand how they define "validated" in this process; for an engineer like me, it means the final output needs to be usable and maintainable by production teams.
Lalam: The system must prove its success through evidence, which is a powerful concept that ensures the outcome is trustworthy before the process is considered done.
Tom: And Jane’s point about validating everything suggests they aren't just trusting the agent; they're forcing it to provide machine-checkable proof of its own work at every stage.
Jane: It feels like a closed loop where every step generates evidence that must be checked before the next step is allowed to happen, which makes sense for robust cloud management.
Lu: The authors are essentially defining a clear path from abstract intention to concrete, verified deployment artifact using this transformation mechanism.
Meng: That’s the core deliverable: taking a vague business request and turning it into an auditable piece of software infrastructure that is guaranteed to function as intended.
Lalam: It means that the system is not just executing commands; it's fulfilling a promise backed by verifiable operational proof, which is a huge cultural shift in how we trust software.
Tom: That clear transition from natural language to verified output is what makes this entire concept feel so powerful for modern cloud operations.
Improvements: Jane: Moving beyond the summary, we need to understand how this framework improves upon existing solutions, particularly regarding safety and reliability when we use AI agents in critical systems.
Tom: The introduction of a zero-trust agent harness is the biggest differentiator, which enforces strict boundaries on what an agent can actually do or even access.
Lu: I find the focus on "Zero-Trust" fascinating because it means we are applying rigorous security principles to the agents themselves, not just to the infrastructure they happen to touch.
Meng: That's exactly what I need, Lu; if an agent is tasked with deployment, it should only have access to those specific APIs and nothing more than that.
Lalam: The idea of bounded recovery also addresses a major cultural issue where AI failure was often treated as an unrecoverable disaster that we just accepted.
Jane: It’s not just failure; the system has a strategy for bounded diagnosis and repair, which is so much more sophisticated than simply letting the agent crash.
Tom: And that's tied into "Evidence-Gated Progression," which prevents us from blindly trusting an agent's report of success without objective proof that we can check ourselves.
Lu: The framework forces a pause and a verification step at every juncture, Lu, ensuring that we are not just accepting a model's assertion but verifying the actual state of the system.
Meng: From an operational standpoint, this mechanism means I can trust the workflow to stop and fix itself when things go wrong within defined limits.
Lalam: It provides confidence in automated systems by providing a safety net that ensures failure isn't just accepted but is actively diagnosed and corrected within the bounds of a plan.
Tom: So, the combination of strict policy enforcement through the harness and built-in recovery loops is what makes this a much more robust system than what we have today.
Conclusion: Jane: We’ve covered so much ground—from the structure to the specific safety features—that we need to look at how this paper puts all its parts together into a cohesive, powerful system.
Tom: It's clear that this architecture is not just theoretical; it has been rigorously tested and proven reliable across multiple models in their experiments.
Lu: I think the biggest implication is that we are finally moving toward scalable, truly autonomous cloud systems, Lu, where the AI can handle complexity itself.
Meng: The practical impact for my startup is a reduction in manual oversight; the system handles complex failures independently, which saves massive amounts of operational time.
Lalam: It elevates our relationship with AI from a simple automation tool to an integrated partner that enforces organizational principles and accountability across processes.
Tom: We’ve seen how it works on Google Cloud, but the framework is designed to be platform-independent, which is another huge benefit for almost any cloud deployment.
Jane: It’s great that this allows the systems to handle complex tasks like data ops or MLOps with the same level of reliable rigor as a human engineer.
Lu: The consistency across those fourteen domains suggests that this truly universal architecture can apply everywhere, Lu, regardless of the specific cloud service needed.
Meng: I'm looking forward to seeing how we implement bounded recovery in real-world scenarios, Meng, since the data shows it works even under stress.
Lalam: It offers a clear path toward operational excellence and ensuring the systems are both efficient and ethically sound in their operations.
Final Conclusion: Tom: We’ve been looking at how this architecture works, but we need to wrap up by summarizing what this paper, "Towards Agentic Cloud Engineering: Graph and Loop Engineering with a Zero-Trust Agent Harness," really means for the industry.
Jane: It really means that we are moving beyond simple automation and achieving reliable autonomy, Jane. The system is designed not just to get the job done but to prove it through verifiable evidence at every single step.
Lu: I think the sheer scope of what this allows—the ability to handle complex, long-horizon tasks across all those fourteen domains—is truly staggering when we consider the possibilities for creative AI workflows.
Meng: I'm excited that it provides such a robust framework for operationalizing these boundaries; it solves the problem of having agents act autonomously but without losing control.
Lalam: This capability is a huge step toward operational excellence, Lalam, ensuring our commitment to rigor is met by empowering AI.
Tom: I agree with Lalam; it’s not just about efficiency, it' about establishing trust in the system itself as a core operational value.
Jane: That's right; the framework guarantees that even if things go wrong, the recovery process is bounded and auditable, which provides immense peace of mind for managers.
Lu: It’s like we are creating a self-healing cloud platform guided by verifiable logic rather than just hoping that the agent works correctly.
Meng: That's exactly the practical impact—the ability to manage failure as a planned part of the repair process, Meng.
Lalam: And I think that gives us a new standard for how autonomous systems should behave in the world going forward.
Tom: Well, that wraps up our discussion of this incredible research paper. Thank you all for joining us today on "Towards Agentic Cloud Engineering: Graph and Loop Engineering with a Zero-Trust Agent Harness."
Jane: We'll be back with another fascinating paper very soon, so make sure to tune in next time!
Sagar Srinivas Sakhinana, Venkataramana Runkana
Tata Research Development and Design Centre
cs.SE, cs.AI, cs.LG
Submitted: 2026-08-30
Updated: 2026-08-30
Comments: Nil
License: http://creativecommons.org/licenses/by-nc-sa/4.0/
Importance score: 81/100
The gist: Agentic Cloud Workflow Engineering provides a comprehensive framework for synthesizing, deploying, verifying, and repairing complex cloud-engineering implementations across diverse domains such as
Key concepts
- Zero-Trust Agent Harness
- This is a key differentiator that enforces strict boundaries on AI agents. It ensures an agent only has access to specific APIs needed for its task, applying rigorous security principles to the agents themselves rather than just the infrastructure.
- Evidence-Gated Progression
- This mechanism prevents blind trust in an AI agent's report of success. It forces a pause and a verification step at every stage, requiring objective proof that the system must be checked before proceeding to the next step.
- Bounded Recovery
- This provides a strategy for handling AI failure. Instead of accepting crashes, the system has built-in mechanisms to diagnose and repair issues within defined limits, ensuring failure is actively corrected rather than ignored.
Terminology
Summary
Agentic Cloud Workflow Engineering provides a comprehensive framework for synthesizing, deploying, verifying, and repairing complex cloud-engineering implementations across diverse domains such as DevOps, CloudOps, SecOps, and MLOps. This work advances the field by establishing a structured methodology that governs the entire lifecycle of an agent's interaction with cloud infrastructure. The core advancement lies in composing three complementary control mechanisms—graph engineering, loop engineering, and agent harness engineering—into a unified architecture that ensures execution is both goal-directed and rigorously verifiable under zero-trust principles.
Graph Engineering
This abstraction denotes the explicit organization of an agentic workflow into a formal structure. It determines where execution may progress
by defining nodes, transitions, branches, dependencies, approvals, recovery paths, and terminal states. Graph engineering establishes the admissible execution paths and the conditions governing progression, recovery, and termination.
Workflow nodes can represent agents or actions while edges encode the necessary handoffs and conditional control flow. For instance, a passing test permits forward progression along a defined path; conversely, failure activates an explicit diagnosis–repair path before testing can be repeated.
Loop Engineering
Loop engineering governs the bounded iterative process required for robust system adaptation when initial attempts fail or require refinement. It specifies how execution responds to feedback and failure.
This mechanism manages the cycle where an agent acts, observes execution or evaluation feedback, diagnoses failures, revises its actions, and retries. The process is strictly bounded:
-
It specifies when recovery is triggered and how failure evidence informs diagnosis and repair or re-planning.
-
It determines whether another attempt is permitted and when recovery must terminate.
Failure evidence drives the system through a cycle of Diagnose to Repair / Re-plan to Retry,
ensuring that the process remains within defined Recovery Bounds.
Agent Harness Engineering
This component defines the controlled runtime and policy boundary that mediates all interactions between the foundation model and external systems. The harness acts as a crucial mediation layer, governing capabilities, permissions, and execution environments in a zero-trust manner. It dictates:
-
Tool routing and access control.
-
The enforcement of
Policy Approval Gates.
-
The maintenance of context and state across extended tasks.
Ultimately, the harness determines which capabilities may be invoked, which resources are accessible,
ensuring that all actions are authorized, isolated, observed, and audited before external execution occurs.
Unified Graph–Loop–Harness Architecture
These three concerns do not operate as sequential stages but compose a single execution framework. The system begins by binding the natural-language task to an identity, tenant, scope, approvals, and policy context. The composition governs agent coordination through the following control surfaces:
-
Graph Engineering: Manages workflow orchestration and
evidence-gated progression.
-
Loop Engineering: Manages bounded recovery through diagnosis and re-verification.
-
Agent Harness Engineering: Enforces authorization, providing
scoped capabilities
and isolation safeguards.
The overall workflow state, telemetry, audit records, and machine-checkable verification evidence provide the control substrate for progression. The system is designed to terminate either in verified completion
or in an auditable failure state
once the required verification conditions remain unsatisfied after the applicable recovery bounds are exhausted.
Improvements for AI systems
The primary improvement is the formalization and mandatory integration of a Unified Graph–Loop–Harness Architecture into agentic workflow design. Current AI systems often treat these components sequentially or ad-hoc; the improvement mandates their simultaneous, compositionally controlled execution to guarantee safety and verifiable progress.
- Mandatory Implementation of Graph Engineering for Workflow Admissibility:
-
Improvement: The system must transition from merely following a sequence of prompts/tools to explicitly defining the workflow as a graph structure (nodes, edges, dependencies). Progression between any two steps must be governed by machine-checkable evidence-gated transitions.
-
Capability Gained: The AI can no longer execute an action simply because it thinks it should. It must prove that the current state satisfies the pre-conditions (the
admissible path
) required for the next transition, eliminating invalid or unauthorized workflow jumps.
- Integration of Bounded Loop Engineering for Guaranteed Recovery:
-
Improvement: Failure handling must be formalized into a bounded, iterative loop structure rather than simple retries. This requires systematic diagnosis (failure evidence to diagnosis to re-planning). The system must strictly enforce recovery bounds (maximum retries, time limits, cost budgets) and terminate in an auditable failure state if those bounds are exhausted.
-
Capability Gained: The system achieves guaranteed termination. It cannot get stuck in infinite loops or spend unbounded resources on failed tasks. It provides predictable failure modes crucial for high-stakes environments (e.g., financial, infrastructure).
- Enforcement of Agent Harness Engineering for Zero-Trust Execution:
-
Improvement: All external tool calls and model invocations must pass through a comprehensive, policy-aware runtime mediation layer (the
Harness
). This layer must enforce granular permissions, manage state isolation (sandboxing), and require explicit approval gates for sensitive operations. -
Capability Gained: The AI achieves policy-scoped capability use. It cannot exploit vulnerabilities by calling unpermitted APIs or accessing resources outside its defined scope, drastically mitigating risks like privilege escalation or cross-tenant data leakage.
The resulting system is a Verifiable, Resilient Cloud Workflow Executor capable of:
-
End-to-End Synthesis and Remediation Across Heterogeneous Domains: It can ingest a high-level natural language specification (e.g.,
Migrate the payment service to AWS EKS and ensure compliance with PCI DSS
) and execute the entire lifecycle—from generating Terraform/Kubernetes manifests (CloudOps), to running security scans (SecOps), deploying, verifying, and automatically repairing any failed component (SRE/AIOps)—all while maintaining a verifiable audit trail. -
Resilient Operation Under Real-World Failure Conditions: It can autonomously diagnose cascading failures (e.g., network degradation affecting a database connection) by entering the bounded loop, systematically re-planning the fix, and re-verifying the entire sequence until either success or definitive failure is reached, all without human intervention beyond initial scope definition.
-
Guaranteed Compliance and Auditability: Every single transition—from initial plan generation to final resource modification—is recorded with machine-checkable evidence (e.g.,
Security Policy X was verified against Evidence Y
). This allows the system to provide an immutable, auditable proof of compliance for every action taken, which is critical for regulated industries.
Abstract
Agentic AI is enabling cloud-based workflows in which autonomous agents reason over operational state, invoke authorized tools, modify software and infrastructure, deploy services, verify execution outcomes, and adapt across long-horizon, multistep tasks. Engineering such workflows requires explicit mechanisms for workflow progression, constrained execution, failure recovery, and verifiable completion. We present Agentic Cloud Workflow Engineering, an agentic AI framework that transforms natural-language agentic cloud-engineering tasks into validated code repositories and verified operational cloud deployments for automating cloud-based agentic workflows. The framework separates three complementary concerns: graph engineering specifies long-horizon workflow progression and verification-dependent transitions; loop engineering provides bounded diagnosis, repair or re-planning, retry, and re-verification; and agent harness engineering enforces zero-trust execution through identity, authorization, policy-scoped capabilities, isolation, and runtime safeguards. Workflow progression and completion require machine-checkable repository, deployment, and runtime evidence, with recovery constrained by explicit operational bounds and termination criteria. We instantiate the framework on Google Cloud and evaluate repository completeness, controlled execution, evidence-gated progression, operational deployment, and bounded recovery. Experimental results show that executions terminate with either a verified operational cloud deployment or an auditable terminal failure under bounded recovery. The framework provides a unified engineering architecture for cloud-based workflows spanning Agentic DevOps, Agentic CloudOps, Agentic SRE/AIOps, Agentic SecOps, Agentic DataOps, Agentic MLOps/LLMOps, AgentOps, Agentic RAG/GraphRAG, and related cloud-engineering domains.
Sources
- Large Language Models for Agentic NetOps and AIOps: Architectures, Evaluation, and Safety
- AgentOps: Enabling Observability of LLM Agents
- From Local to Global: A Graph RAG Approach to Query-Focused Summarization
- SAGA: Workflow-Atomic Scheduling for AI Agent Inference on GPU Clusters
- MLOps Monitoring at Scale for Digital Platforms
- ARPaCCino: An Agentic-RAG for Policy as Code Compliance
- kRAIG: A Natural Language-Driven Agent for Automated DataOps Pipeline Generation
- Authenticated Delegation and Authorized AI Agents
Related papers
- Falsification-Based Verification of LLM-Generated Optimization Models: Sound Test Batteries and Their Detection Limits
- GitSkills: A Dataset of Agent Skills on GitHub
- SABER: Benchmarking Operational Safety of LLM Coding Agents in Stateful Project Workspaces
- PackMonitor: Enabling Zero Package Hallucinations Through Decoding-Time Monitoring
- IntentCoding: Amplifying User Intent in Code Generation
- Incentives and Outcomes in Bug Bounties