Emergence-as-Code as a Foundation for Reliable Self-Governance
Listen
Radio episode about this paper
Transcript
Introduction to the show: ident: Robotics Radio. Generated commentary on the latest robotics and control papers.
Rosa: I'm Rosa, and with me are Dev and Taro, guest researcher.
Dev: Today's paper: "Emergence-as-Code as a Foundation for Reliable Self-Governance".
Rosa: Service-level objective (SLO)-as-code tools make per-service reliability declarative, but users experience journeys: end-to-end executions whose availability and tail latency emerge from topology, routing, redundancy, timeouts/fallbacks, shared failure domains,
Dev: First, who's behind it and why it matters.
Title and authors: Rosa: Moving on to a deeper look at the actual content of "Emergence-as-Code as a Foundation for Reliable Self-Governance," the paper summarizes how EmaC provides this declarative contract for journey-level SLO bounds and governance artifacts from intent and evidence.
Dev: They are essentially saying that traditional SLO-as-code tools often manage services in isolation, but that end-to-end executions have reliability characteristics that emerge from topology, routing, redundancy, and all those factors we usually ignore.
Taro: The summary emphasizes that the core innovation is creating a declarative contract—the EmaC specification—that defines how intent is fed into Model Discovery to generate evidence-backed deltas for edges and failure domains.
Rosa: And this means the system doesn't just look at what we want; it actively proposes potential changes to the system structure based on what it sees in traces, mesh data, or deployment metadata.
Dev: This process culminates in a compiler that derives those optimistic and pessimistic journey bounds, which are critical because they quantify both the best-case and worst-case scenarios for availability and latency.
Taro: It seems the paper is laying out a structured way to handle uncertainty by explicitly calculating these bounds based on failure-domain assignments, which is a very concrete way to deal with ambiguity in complex systems.
Rosa: And they define the semantics using specific operator definitions, like how Parallel or Race operators evaluate availability and latency under those different conditions.
Dev: I read about the "Optimistic bound" assuming independence where permitted in the operator composition sketch, which is a direct way to model an ideal scenario based on certain assumptions.
Taro: That’s interesting because they also have a "Pessimistic bound" that explicitly collapses redundancy inside an effective failure domain, which models how real-world constraints limit gains.
Rosa: So the summary highlights that the contract is designed to be executable, and it produces three artifact classes: synthetic SLIs for availability, alerting and release artifacts like burn-rate rules, and a provenance report.
Dev: The provenance report is what gives accountability because it links every generated rule directly to the operator expression, evidence window, confidence threshold, and failure-domain assumption that produced it.
Taro: That level of traceability is essential for trust; if something goes wrong later, you need to know exactly what inputs led to that specific decision.
Rosa: Ultimately, the summary paints EmaC as a mechanism that takes declarative SLOs and turns them into actionable, bounded governance artifacts through an iterative cycle of model discovery and compilation.
Dev: It shifts the focus from simply declaring local service reliability to managing the emergent properties of the entire journey reliably across a distributed system.
The paper's summary: Rosa: Now let’s look at what the paper suggests as improvements to this concept, focusing on how EmaC can be made more effective in practice, which is where we can really apply this idea.
Dev: One major improvement is shifting from reactive management to a proactive system where the AI evolves into an automated EmaC compiler and controller managing journey reliability.
Taro: That means instead of waiting for something to break, the AI should be continuously running Model Discovery to propose versioned, evidence-backed changes to system topology and failure domain assumptions proactively.
Rosa: This leads directly into a more precise decision gate for deployment where we can evaluate complex guards, like checking if the pessimistic journey availability bound is above zero point nine nine five alongside latency targets and confidence scores.
Dev: That’s critical because it allows us to move beyond simple local checks and evaluate the actual risk associated with shared-fate rules before promoting something to production.
Taro: I also see this being used for continuous drift detection across distributed systems, automatically detecting when topology changes, which would immediately trigger a recalculation of the pessimistic bound if the system shifts.
Rosa: And then they suggest generating that auditable governance artifact, where every accepted model configuration is linked to its provenance report detailing exactly which evidence and policy decisions led to it.
Dev: Furthermore, there’s the idea that this AI can mechanically compose complex latency metrics from raw Prometheus histograms and traces into synthetic SLIs using techniques like convolution or mixtures.
Taro: That way we aren't ignoring tail correlations; we are modeling them directly instead of just ignoring them in favor of simpler averages.
Rosa: And finally, they suggest making the uncertainty explicit by treating dependence assumptions about independence or redundancy as compiler-visible properties that are continuously reconciled from evidence and used as dynamic action guards.
Dev: So essentially, the improvement is to embed uncertainty directly into the control loop so that automation only acts when evidence and bounds agree with those explicit assumptions.
The paper's improvements: Rosa: To wrap things up on "Emergence-as-Code as a Foundation for Reliable Self-Governance," the paper proposes this minimal way to govern the gap between local SLO-as-code and emergent journey reliability by treating the model as a reviewable hypothesis and making every generated decision carry its assumptions.
Dev: It seems like they are achieving this by creating a self-governing system where automation only proceeds when evidence, bounds, and policy all align, otherwise surfacing the uncertainty as a diff.
Taro: I think the main implication is that we get better at managing complex systems by making the uncertainty explicit instead of hiding it in implicit behavior.
Rosa: This approach keeps self-governance compatible with existing service ownership while providing a path toward systems where automation only acts when evidence, bounds, and policy agree.
Dev: So it’s about creating a system that is auditable at every step, ensuring that we know precisely why a decision was made.
Taro: I think the paper opens up avenues for how we can build more resilient production environments by focusing on explicit risk assessment across the entire journey instead of just local checks.
Rosa: So, "Emergence-as-Code as a Foundation for Reliable Self-Governance" gives us a framework to manage that tricky gap between local SLOs and system behavior.
Dev: It’s about making sure that self-governance is tied directly to verifiable evidence and policy decisions.
Taro: By making the uncertainty part of the model, we get a way to handle complexity without having to hide the unknowns.
Conclusion: Rosa: So, to wrap up, "Emergence-as-Code as a Foundation for Reliable Self-Governance" essentially shows how we can move beyond isolated service reliability and start governing entire user journeys declaratively through this contract approach.
Dev: That’s right, it's about turning intent and evidence into concrete journey bounds that we can actually check against our deployment policies.
Taro: I think the real impact here is moving away from reactive fixes toward a proactive system that understands how the whole architecture behaves under stress or when things go wrong.
Rosa: It’s fascinating to think about applying this outside of controlled lab environments, though I’m curious how long you reckon this kind of self-governing loop can sustain itself in a truly wild, unpredictable field?
Dev: The loop rate is key here; if the discovery and compilation cycle takes too long, we're just reacting to old evidence by the time we get a new bound. I wonder about the latency involved in that whole process.
Taro: That’s where Model Discovery comes in, proposing deltas based on evidence records, and if that inference takes too much time, the autonomy of the system suffers. It has to be fast enough to keep up with real-world misbehavior.
Rosa: And speaking of real-world misbehavior, how robust is this contract when we hit those shared failure domains we talked about? Does it truly model the pessimism correctly under extreme conditions?
Dev: The pessimistic bound collapses redundancy inside an effective domain, which should give us a very conservative view of the availability. It means if things are sharing a fate, we’ll see that bound drop significantly.
Taro: That conservatism is what makes it useful; it forces automation to be very cautious when evidence is weak or when there’s ambiguity about independence assumptions. It stops the system from blindly trusting optimistic assumptions.
Rosa: I think this level of accountability, with that provenance report linking every decision back to its evidence and assumptions, really builds trust in these self-governing systems.
Dev: Absolutely; that audit trail is crucial for debugging when things go sideways in a complex distributed environment, showing exactly which policy and evidence led to the accepted model.
Taro: It sets a high bar for how we define autonomous control in production settings, forcing us to be explicit about what we assume versus what the data actually shows.
Rosa: So it seems like this paper gives us a solid foundation for building systems that don't just run, but that actively manage their own reliability under real-world pressure.
Dev: Indeed; "Emergence-as-Code as a Foundation for Reliable Self-Governance" gives us the tools to bridge that gap between local SLOs and emergent system behavior.
Taro: Next up, we’ll see how these models integrate with the temporal network studies, which is going to be interesting for understanding control in dynamic systems.
Innopolis University
cs.SE, cs.DC, cs.PF, cs.SY, eess.SY
Submitted: 2026-02-05
Updated: 2026-10-01
Code: https://github.com/a-a-k/Emergence-as-Code
License: http://creativecommons.org/licenses/by-sa/4.0/
Importance score: 90/100
The gist: Service-level objective (SLO)-as-code tools make per-service reliability declarative, but users experience journeys: end-to-end executions whose availability and tail latency emerge from topology,
Key concepts
- Intent vs. Evidence
- Intent is what the user declares should hold (the goal). Evidence is the actual data observed from the system (what it appears to be doing). EmaC separates these two, using evidence to propose a model that bridges the gap between the desired intent and real-world behavior.
- Model Discovery
- This process proposes a versioned hypothesis about how the journey works based on collected evidence. It suggests specific details like redundancy groups or failure domains, providing provenance and confidence scores for each suggestion so reviewers can assess its reliability.
- Optimistic vs. Pessimistic Bounds
- EmaC calculates two bounds: optimistic assumes components are independent (best-case scenario), while pessimistic removes unvalidated independence assumptions, collapsing redundancy within a single failure domain to create a safer, more conservative estimate of the journey's true performance.
Terminology
Summary
Service-level objective (SLO)-as-code tools make per-service reliability declarative, but users experience journeys: end-to-end executions whose availability and tail latency emerge from topology, routing, redundancy, timeouts/fallbacks, shared failure domains, and tail amplification. This paper proposes Emergence-as-Code (EmaC), a declarative contract that compiles journey-level SLO bounds and governance artifacts for declared SLOs from intent and evidence.
How it works
EmaC separates intent
(what should hold) from evidence
(what the system appears to be doing). The core control loop involves versioned intent feeding Model Discovery, which proposes an evidence-backed journey model, followed by a compiler that derives bounded SLOs and governance artifacts. This process is iterative: runtime evidence continuously triggers Model Discovery and proposed deltas.
The EmaC contract defines a typed journey expression, leaf bindings to atomic SLOs and telemetry, failure-domain assumptions, and guarded actions. The compiler derives optimistic and pessimistic journey bounds
based on these inputs. For example, the optimistic availability assumes independence where permitted in the operator composition sketch: Parallel(J1,..., Jn) A+ = Î i Ai; L = max i Li for join-all fan-out.
Model Discovery and Evidence Synthesis
Model Discovery proposes evidence-backed deltas for edges, branch probabilities, redundancy groups, and failure-domain hypotheses,
with each delta carrying provenance and confidence.
The discovered model is a versioned hypothesis whose confidence and provenance are visible to reviewers and controllers.
This discovery process uses an EvidenceRecord
(normalized observation from traces, mesh/traffic policy, deployment metadata, or operator declaration) to create a CandidateModel
(an operator graph plus inferred attributes).
Compositional Semantics and Bounded Bounds
EmaC utilizes a small operator algebra where each journey expression evaluates to availability A and latency random variable L. Given a failure-domain assignment D, EmaC computes an interval [A−, A+]. The optimistic bound assumes independence, while the pessimistic bound removes unvalidated independence assumptions within each effective domain: redundancy within one domain provides no gain.
This results in a signal that the journey depends on unvalidated independence assumptions.
Replay and Governance Artifacts
To ensure executability, EmaC includes a deterministic checkout replay. This replay consumes the intent, synthetic trace evidence, and policy threshold to emit discovered models and typed deltas. The compiler emits three artifact classes: synthetic SLIs (pessimistic availability, optimistic availability), alerting/release artifacts (burn-rate rules, progressive-delivery gates), and a provenance report linking each generated rule to the operator expression, evidence window, confidence threshold, and failure-domain assumption that produced it.
Enforcement and Decision Making
The compiler generates a guard
for promotion: "allow(promote) iff A− >=.995 && p99 U =.8 && no unresolved domain conflict. Low-confidence deltas are not auto-applied; they become
reviewable diffs, ensuring that
automation defaults to conservative decisions under shared fate or weak evidence. This mechanism makes the reconciliation process
auditable: every accepted model is linked to the evidence records and policy decision that admitted it."
Assumptions and Boundaries
The proposal makes several explicit assumptions to maintain safety. First, Model Discovery is not assumed to recover the true architecture but proposes a typed hypothesis about the effective journey model under a stated evidence window.
Second, availability semantics use dependence extremes: optimistic assuming independence where permitted, and pessimistic collapsing redundancy inside an effective failure domain. Third, Race and Timeout assume cancelable alternatives in the basic semantics,
treating production behaviors like retries as evidence-backed extensions rather than implicit behavior.
The contract ensures that Conservative default allows candidate bounds for analysis, but forbids promotion from an inferred independence assumption until the model is accepted; hence weak evidence yields REVIEW.
Conclusion
EmaC proposes a minimal way to govern the gap between local SLO-as-code and emergent journey reliability by treating the model as a reviewable hypothesis and making every generated decision carry its assumptions.
This approach keeps self-governance compatible with existing service ownership while providing a path toward self-governing reliable systems: automate only when evidence, bounds, and policy agree; otherwise surface the uncertainty as a diff.
The gist
EmaC compiles journey-level SLO bounds and governance artifacts for declared SLOs from intent and evidence.
Table 1: EmaC contract: what is declared, inferred, and checked.
Improvements for AI systems
Based on the provided research paper, here are specific improvements that can be made to AI systems, categorized by the capability they would gain:
The core improvement is shifting from reactive, local SLO management to a proactive, evidence-backed system for self-governance across complex journeys.
-
The AI system can evolve into an automated
Emergence-as-Code
(EmaC) compiler and controller that manages the reliability of end-to-end user journeys rather than individual services in isolation. -
The improved AI can ingest declarative intent (the desired journey objective) and real-time telemetry to continuously run a complex model discovery engine, proposing versioned, evidence-backed changes to the system's topology and failure domain assumptions.
-
The AI can derive
compiled
journey SLO bounds (optimistic and pessimistic) that explicitly quantify the risk associated with current system state (e.g., shared failure domains or redundant components).
Specific functional improvements enabled by these capabilities:
-
The system will provide a precise, actionable decision gate for deployment: Instead of simply checking if a service's local SLO is met, the AI will evaluate a complex guard like: "Allow promotion only if the pessimistic journey availability bound (derived from shared-fate rules) is greater than 0.995 AND the p99 latency bound is under 400ms AND evidence-backed confidence in the current failure domain model exceeds 80%."
-
The system will perform continuous
drift detection
across distributed systems: It will automatically detect when system topology changes (e.g., two previously independent payment providers are now sharing a single effective failure domain) and immediately trigger a recalculation of the pessimistic journey bound, potentially changing the deployment decision fromPass
toFail/Review.
-
The system will generate an auditable governance artifact: For every accepted model configuration, it will emit a provenance report detailing exactly which piece of evidence (trace spans, mesh routes) and policy decision (confidence threshold) led to the final bound and action guard. This allows for complete accountability in self-governing systems.
-
The system can mechanically compose complex latency metrics: It will ingest raw Prometheus histograms and OpenTelemetry traces to mechanically compose discretized histograms using convolution, mixtures, and order statistics to emit synthetic SLIs consumed by alerting systems, ensuring that tail correlations are modeled rather than ignored.
-
The AI will manage uncertainty explicitly: Instead of hiding assumptions about independence or redundancy within a failure domain, it will make these dependence assumptions a
compiler-visible property
of the system model, continuously reconciled from evidence and using the resulting bound as a dynamic action guard.
Abstract
Local adaptations can preserve component health while changing whether a system meets its reliability requirements. The system-level consequence depends on the interaction affected and its role in the user journey. Emergence-as-Code (EmaC) makes this relationship part of an executable, evidence-reconciled CompositeSLO assessment. A declared journey obligation persists while discovery maintains a hypothesis of its runtime realization. An inferred operator-to-role binding selects the measurements used in the calculation; its evidential status determines whether the binding supports a numerical assessment. Each result retains the obligation and versioned hypothesis as a basis for governance decisions. A PetClinic proof of concept recovered the exact operator-state/edge-binding change in all 20 randomized treatments, with no false change in 20 controls. EmaC and a manual dynamic composite matched disjoint semantic outcomes in all 40 conditions. A model with frozen suppression knowledge overestimated required-history availability by 0.01, placing treatments above a 0.995 target despite observed availability of 0.99. Ambiguous and contradictory evidence produced UNASSESSABLE. The research programme evaluates sustained binding maintenance, calibrated uncertainty, and the decision value of this requirement-level account of local adaptation.
Sources
- Evaluating Asynchronous Semantics in Trace-Discovered Resilience Models: A Case Study on the OpenTelemetry Demo
- Principles of Antifragile Software
Related papers
- Falsification-Based Verification of LLM-Generated Optimization Models: Sound Test Batteries and Their Detection Limits
- GitSkills: A Dataset of Agent Skills on GitHub
- SABER: Benchmarking Operational Safety of LLM Coding Agents in Stateful Project Workspaces
- PackMonitor: Enabling Zero Package Hallucinations Through Decoding-Time Monitoring
- IntentCoding: Amplifying User Intent in Code Generation
- Incentives and Outcomes in Bug Bounties