Steelhead: Interleaving Partially Synchronous and Asynchronous Commit Rules on a Shared DAG

arXiv:2609.30163 · cs.DC, cs.CR · Submitted 2026-09-24 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: Security Radio. Generated commentary on the latest security and cryptography papers.

Nadia: Today's paper: "Steelhead: Interleaving Partially Synchronous and Asynchronous Commit Rules on a Shared DAG".

Elias: Steelhead is a dual-mode consensus protocol that composes a partially synchronous and an asynchronous commit rule over one DAG: "every k-th round is decided by the asynchronous rule,

Nadia: First, who's behind it and why it matters.

Title and authors: Nadia: So, we're diving into "Steelhead: Interleaving Partially Synchronous and Asynchronous Commit Rules on a Shared DAG." It sounds like this paper is tackling the problem of making consensus protocols that can handle both stable, fast networks and those where things are really messy.

Elias: Exactly; the title tells us they're combining two different ways of agreeing—a partially synchronous rule with an asynchronous one—onto a single Directed Acyclic Graph. It seems like they’re trying to get the best of both worlds without needing separate systems for each scenario.

Priya: From my angle, I'm curious about what this combination actually means for privacy or measurement, since we often deal with data that needs strict ordering. If you can interleave these rules, does that mean the resulting protocol has a predictable commitment structure regardless of whether the network is behaving nicely or randomly?

Nadia: That’s a huge question, Priya. What they're pointing out is that they don't need to guess beforehand which rule to use; instead, they let the committed DAG tell them how to decide each round based on the round number.

Elias: And that decision logic hinges on this wavelength function w(r), which changes depending on whether the round number is divisible by some period k.

Nadia: Right, Elias. It’s not a simple switch; it's a dynamic decision process where the protocol adapts to the current state of the DAG rather than relying on external timing mechanisms.

Priya: I wonder if this adaptation means that even in very noisy environments, there’s still some guarantee about how long we wait for finality?

Elias: That’s where they get clever with their period adaptation mechanism, Algorithm two. It explicitly states that the pivot for changing the period can't come from the ledger itself because under asynchrony, it stalls below the first synchronous slot left undecided.

Nadia: So, if things get really slow and asynchronous, this mechanism prevents the protocol from locking into a bad period by keeping it from falling too far.

Priya: That sounds like a safety net against getting stuck in an unproductive state during periods of high network latency or instability. Does this adaptability translate into better data integrity for applications?

Elias: It does, because the paper claims that the asynchronous rule applied to those coin rounds alone keeps the protocol live, even when things are very slow.

Nadia: That's powerful because it means liveness isn't completely sacrificed just to handle asynchrony; they keep every slot on track by having a hidden leader for them.

Title and authors: Priya: So, if we think about the data flow, this suggests that the protocol can maintain a consistent commitment structure even when the network is highly variable. It’s like having two different ways to organize a filing system that automatically shifts based on how chaotic the mail delivery is.

Nadia: Precisely, Priya. And looking at their evaluation metrics, they show that Steelhead matches the latency of both the partially synchronous and asynchronous protocols under specific conditions—claims C1 through C4 hold even at n = fifty.

Elias: The results are quite compelling because they track the better protocol within a few percent across several difficult network conditions.

Priya: Tracking performance against multiple delay scenarios suggests that this isn't just theoretical work; it’s showing practical resilience when we consider real-world network jitter and random delays. It moves the discussion away from ideal models toward how these systems perform under pressure.

Nadia: Absolutely, Priya, and that robustness is what makes this interesting for applied security research. We need to know how cheap it is to break this system, right?

Elias: That's a crucial question because if the underlying assumptions about the coin’s unpredictability or the floor of committed candidates are violated—which they explicitly state aren't derived from their model—then we open up avenues for attack.

Priya: So, when you look at those technical assumptions, Elias, what seems like the weakest link from a data integrity standpoint? Where does the theoretical guarantee start to rely on something that might be fragile?

Elias: The paper highlights that they haven't derived the unpredictability of the coin or the per-round floor of committed candidates from clause A5; they treat those as hypotheses.

Nadia: That’s a big caveat for us; it means our security analysis has to start by assuming those properties hold true, which is a necessary starting point, but it also means we have to rigorously check those assumptions ourselves.

Priya: And that brings us back to the data itself. Since the model doesn't execute and relies on hypotheses, how much of what we’re seeing in these performance results is based on the abstract structure versus actual hardware or message scheduling?

Nadia: Well, Priya, their evaluation shows they instantiated Steelhead on two pairs of protocols: Mysticeti with MahiMahi and BlueBottle’s variants. This means they've tested the mechanism against existing, established DAG commit rules.

Elias: And by showing how it compares to those specific protocols under heavy load—like a full random delay probability of one—they give us some concrete benchmarks for performance comparisons.

Title and authors: Priya: That gives me something tangible to work with; we can look at the actual message schedules and see if their performance claims hold up when we measure things in the real world, not just on paper.

Nadia: Exactly, Priya. And that’s where the real impact could be; if this structure can indeed match those latencies across conditions, it means protocols built on this dual-mode approach could be far more reliable for distributed systems that operate in unpredictable environments.

Elias: It suggests a path forward where we don't have to choose between a protocol optimized for speed in the best case and one optimized for liveness in the worst case; Steelhead tries to bridge that gap using round-number dependent logic.

Priya: It’s an interesting architectural idea, Elias, but I still want to stress that without seeing how this translates into concrete data structures or measurement techniques, it remains a compelling theoretical framework for robust consensus.

Nadia: Well, we've got a solid overview of the "Steelhead: Interleaving Partially Synchronous and Asynchronous Commit Rules on a Shared DAG" paper. It shows us how to blend synchronous and asynchronous rules using round-number dependent logic to maintain consistency on one DAG.

Elias: The key takeaway is the adaptive period selection mechanism, which lets validators recompute the best rule at run time from the committed DAG alone.

Priya: I think for applications, it means we can design systems that are inherently more resilient to network degradation by having a built-in mechanism to switch between operational modes automatically.

Nadia: That’s the practical implication; moving away from relying solely on external timers or complex handshakes for mode switching.

Elias: And that dynamic adaptation, driven by DAG evidence, is what really separates this from other static dual-mode approaches we've seen in the literature.

Priya: So, as we wrap up this discussion on Steelhead, it seems like a very solid piece of theoretical engineering that addresses a real tension in distributed systems design.

Nadia: It really does; it gives us a framework to think about how to build consensus engines that are inherently more flexible when the network environment shifts unexpectedly.

Elias: We should definitely keep an eye on how their formalization handles those assumptions we discussed, because that's where the real security work will live.

Priya: I agree; it’s a strong foundation for future research into making distributed systems that operate reliably across vastly different network conditions.

The paper's summary: Nadia: So, to recap, Steelhead is this dual-mode consensus protocol that manages to combine a partially synchronous rule and an asynchronous rule all over one shared Directed Acyclic Graph using round numbers as the primary decision factor.

Elias: Right, it’s about letting the structure of the DAG dictate whether you’re following a fast, synchronous pace or a slower, more resilient asynchronous pace without needing any external signals or votes to switch between them.

Priya: What this really means is that we get a protocol that can handle both high-throughput environments and those with extreme network jitter simultaneously, which is something most traditional single-mode protocols struggle with.

Nadia: Exactly, Priya; it’s about achieving performance parity across a wider range of network conditions than either pure synchronous or pure asynchronous approaches could achieve on their own.

Elias: And the real magic here, as I see it from a cryptographic angle, is this adaptive period selection mechanism that allows validators to essentially recompute the best operating rule at run time using only the history of committed blocks.

Priya: From a privacy and measurement standpoint, that adaptability means the resulting commitment structure stays consistent even when message delivery times are highly variable or unpredictable.

Nadia: It translates to a system that is inherently more robust against adversarial network behavior, because it doesn't get stuck waiting for a specific timeout or period change signal.

Elias: And the safety arguments they lay out, particularly how the anchor search floor at round r + w(r) ensures indirect decisions align with direct commits, gives us a strong foundation for understanding its integrity guarantees.

Priya: The data shows that even under severe network conditions, like full random delays or high jitter, Steelhead manages to stay within a small margin of error compared to the best single protocol.

Nadia: That’s what we want to hear; it suggests this architecture could be used in real-world distributed systems where you can't perfectly control the network environment.

Elias: And while they lay out some interesting hypotheses about things like coin unpredictability, it’s important to remember those are assumptions they haven't derived from their model, which is a key area for future work.

Priya: So we have a protocol that performs well under stress based on strong structural properties, but the actual security proof still depends on some underlying assumptions we need to investigate further.

Nadia: Exactly; it gives us a great starting point for how to build consensus engines that are inherently more flexible when the network environment shifts unexpectedly.

The paper's improvements: Nadia: We’ve just looked at how Steelhead works, and now we need to talk about what they suggest as improvements for this dual-mode protocol.

Elias: They propose several mechanisms that allow the AI system to function much more flexibly across different network conditions, moving beyond a fixed setup.

Priya: From my side, I’m really interested in how these suggested improvements translate into tangible benefits for data privacy and measurement integrity in real-world scenarios.

Nadia: They suggest an adaptive consensus mechanism that lets the AI system dynamically switch its underlying commit logic based on what it observes on the DAG, instead of relying on preset timers or mode votes.

Elias: That’s significant because it means the system can transition smoothly from a fast, synchronous-like operation when conditions are good to a more resilient asynchronous operation when things get messy.

Priya: That adaptability is what promises better data integrity because the protocol doesn't stall; it just adjusts its commitment pacing to match the current network reality.

Nadia: And then there’s this self-tuning period adaptation, where the system recomputes its operating period using only committed DAG evidence at runtime, which eliminates any need for external configuration or guesswork.

Elias: That deterministic function of window and round numbers they describe is what makes the period change safe; they show that a pivot cannot come from the ledger itself if you're in asynchrony.

Priya: So, it sounds like this mechanism ensures that even during high jitter or random delays, there’s still a predictable commitment structure being maintained across all nodes.

Nadia: Exactly, Priya; the results show that this structural flexibility allows Steelhead to track the better protocol—whether synchronous or asynchronous—across six different network conditions.

Elias: And from an engineering standpoint, their implementation on protocols like Mysticeti and BlueBottle shows how to integrate this wavelength schedule and anchor search floor into existing DAG structures without adding extra messages.

Priya: It’s fascinating how they manage to keep the block format untouched while only changing the round pacing and retention horizon of the DAG layer itself.

Nadia: And that leads us right back to my main concern: exploitation. If this system adapts so well, what's the cheapest way an attacker could potentially try to break these assumptions?

Elias: The authors themselves flag that they haven't derived the coin’s unpredictability or the floor of committed candidates from their model, which means those are still hypotheses we have to verify rigorously.

Priya: So, while it looks structurally sound for high performance under stress, the security hinges on verifying those underlying assumptions about randomness and candidate quotas.

Nadia: Right; so as we move forward, our focus needs to be on testing the limits of those specific hypotheses they've made.

Conclusion: Tom: So we’ve gone through the mechanics of Steelhead, and now Nadia and Elias need to bring us to a close by summarizing why this paper matters for applied security research.

Nadia: Essentially, Steelhead is showing how to build a consensus engine that can operate reliably across vastly different network conditions by intelligently interleaving synchronous and asynchronous commit rules over a single DAG.

Elias: And the core contribution is that it achieves this through an adaptive period selection mechanism, meaning the protocol recomputes its best operating rule at runtime based on committed DAG evidence.

Priya: What this implies for us in privacy and measurement is that we can design systems that are inherently more resilient to network degradation by having a built-in mechanism to switch operational modes automatically.

Nadia: It moves the needle away from relying on external timers or complex handshakes for mode switching, which is a big deal for real-world distributed applications.

Elias: And the safety arguments they provide, based on how anchors are searched across different rules, give us a solid foundation for understanding its integrity guarantees.

Priya: The evaluation data really shows that this approach keeps performance metrics within tight margins even under conditions like high jitter or full random delays.

Nadia: It suggests that this architecture could be used in practical distributed systems where you can't perfectly control the network environment to maintain consistency.

Elias: And while they’ve laid out some interesting hypotheses about things like coin unpredictability, it’s important for us to keep checking those assumptions because that’s where the real security work will live.

Priya: I agree; so we have a protocol that performs well under stress based on strong structural properties, but the actual security proof still depends on verifying those underlying assumptions about randomness and candidate quotas.

Nadia: Exactly; it gives us a framework to think about how to build consensus engines that are inherently more flexible when the network environment shifts unexpectedly.

Elias: We should definitely keep an eye on their formalization details, because understanding those specific hypotheses is what will determine the actual security of this Steelhead protocol.

Priya: And that brings us to where we’ll head next, so let’s see what other interesting papers are on the arXiv today.

Mysten Labs

cs.DC, cs.CR

Submitted: 2026-09-24

Updated: 2026-09-28

Code: https://github.com/gdanezis/lean-dag

License: http://creativecommons.org/licenses/by/4.0/

Importance score: 92/100

The gist: Steelhead is a dual-mode consensus protocol that composes a partially synchronous and an asynchronous commit rule over one DAG: "every k-th round is decided by the asynchronous rule, whose leader a

Key concepts

Steelhead
A dual-mode consensus protocol that combines partially synchronous and asynchronous commit rules onto a single DAG. It uses round numbers to decide which rule to apply, allowing it to handle both fast and messy networks simultaneously without needing separate systems.
Period Adaptation Mechanism
A clever feature where the protocol recomputes its operating period at runtime using only committed DAG evidence. This mechanism ensures that even in asynchronous conditions, the protocol avoids locking into a bad period by preventing it from falling too far below synchronous slots.
Wavelength Function w(r)
The function used to determine which commit rule applies based on the round number (r). This function is dynamic; it changes depending on whether the round number is divisible by some period k, allowing the protocol to adapt its behavior based on the DAG's state.
Hypotheses of Unpredictability
The paper relies on hypotheses regarding things like coin unpredictability and the floor of committed candidates. The hosts note that these properties have not been derived from the model, meaning security analysis must rigorously check these assumptions.

Terminology

Summary

Steelhead is a dual-mode consensus protocol that composes a partially synchronous and an asynchronous commit rule over one DAG: every k-th round is decided by the asynchronous rule, whose leader a common coin reveals after the votes, and all other rounds by the partially synchronous rule.

Key mechanisms include:

- Dual-mode mechanism:

Steelhead sends no message beyond the DAG’s blocks, not even to agree on the period, and opens a coin only on the rounds that need a hidden leader.

DAG commit rules only read the DAG and, beyond pacing, never shape it, so two rules can share one DAG: each round holds one leader slot, and the round number picks which rule decides it.

Every k-th slot is an asynchronous slot, decided by the asynchronous rule (Ra), and the others are synchronous slots, decided by the partially synchronous one (Rs).

There is no mode vote, no path switch, and no extra message.

- Slot Decision Mechanism:

The decision logic depends on a round-number dependent wavelength function: The wavelength function is w(r) = wa if r mod k = 0, and ws otherwise, where k is the period in force at round r (Section 4 indexes it by interval, as kj(r), once it adapts).

Each slot is therefore decided by exactly one rule, fixed by its round number.

The anchor of the slot at round r is searched from round r + w(r), a floor set by the wavelength of the undecided slot, not by that of the anchor.

"This is the core of the safety argument. It lets Steelhead skip the reconciliation step of related work [16,5,11], agreeing on where the previous rule stopped: the slots one rule leaves pending are decided by the anchors above them, whichever rule decides those."

- Period Adaptation Mechanism:

Steelhead adapts its period dynamically based on DAG evidence: The period selects between the two rules from what the committed DAG shows, rather than from a guess about the network, which the adversary controls.

"Algorithm 2 states the loop [for adaptation]: The pivot cannot come from the ledger. Under asynchrony the ledger stalls below the first synchronous slot left ⊥, so a pivot taken from it would never be committed and the period could never fall (Fig. 3)."

The new period is a deterministic function of the pivot’s window and of round numbers, and it applies only from the next interval on.

The replay ranks periods, while the failover guarantees that an output which has stopped advancing always gets the rule that makes it advance again (and greatly simplifies the liveness proof).

- Safety and Liveness:

Safety. A direct commit leaves a certificate in every block from round r + w(r) on, and the anchor search starts there, so an indirect decision by either rule agrees with a direct commit by the other.

"Liveness. After GST, synchronous slots with an honest leader and a leader wait on their successor round commit directly by clause A4, and asynchronous slots do not get in the way (Theorem 2). Under asynchrony, the control reading keeps resolving, because every slot on it has a hidden leader."

- Contributions:

An adaptive period selection mechanism allowing validators to recompute the best commit rule at run time from the committed DAG alone, with no extra message and no timer.

Steelhead, a dual-mode mechanism in which two commit rules interpret one DAG and the round number alone selects which of them decides each slot.

Proofs of safety and liveness, machine-checked in Lean 4.

- Evaluation:

The protocol is evaluated under six network conditions: "a small leader delay of 30 ms, below the timeout; (ii) a large leader delay of 125 ms, above it; (iii) a permanent crash of f validators; (iv) a partial random delay, where each message is delayed with probability 0.3 by a duration drawn uniformly at random in 100–150 ms; (v) a full random delay, the same with probability 1; and (vi) a high jitter, an exponentially distributed delay of mean 75 ms capped at 400 ms."

Claims C1 to C4 hold: "C1: In a healthy network, Steelhead matches the latency of the partially synchronous protocol. C2: Under network conditions that stall the partially synchronous protocol, Steelhead matches the latency of the asynchronous one. C3: Under benign asynchrony, Steelhead matches the latency of the best of the two protocols. C4: Steelhead adapts promptly to the network: it switches to the best protocol when conditions change and back when they lift."

- Implementation:

We implement Steelhead on two pairs of protocols: Mysticeti [3] with MahiMahi [18], at n ≥ 3f + 1, and BlueBottle’s partially synchronous variant with its asynchronous variant [29], at n ≥ 5f + 1.

"Steelhead adds a per-round wavelength schedule read by the decision rule and the anchor search, restricts the leader wait to synchronous slots and canary rounds, and implements the period update by counterfactual replay of Algorithm 3."

"The block format remains untouched, and the DAG layer changes only in its round pacing, which reads the period to select the leader wait, and in its retention horizon: Steelhead adds no protocol messages, no cryptography beyond the coin the asynchronous protocol already carries, and no storage access."

- Formalization:

Every statement of Appendix B is machine-checked in Lean 4.

The model is the DAG and a clock that stamps each block: blocks carry a round, an author and references, a view is a set of blocks closed under reference, and a schedule assigns each slot a round, a leader and a kind.

"What is new for this paper is the decision rule with a per-slot wavelength, the anchor floor at r + w(r), agreement across slot kinds at each validator’s own derived schedule, the control verdicts as Ra on the sub-schedule of one scan’s control slots, the coin as a uniform distribution against an adversary that adapts to every earlier draw, and the handover corollary as the one point where the two rules interact."

The period is a configuration-sequence model written for this paper.

- Results:

Table 8 provides the degraded plateau of the three lines per pair and panel. Steelhead’s time to reach period 1 after the onset and to return to 64 after the condition lifts.

Claims C1 to C4 hold at n = 10 as well.

Steelhead stays within 1% of the partially synchronous protocol (208 vs 206 ms and 163 vs 162 ms) and 27% below the asynchronous one.

"It reaches period 1 within 10–25 s of the onset in every panel where a protocol stalls and returns to 64 within 5–15 s of the condition lifting, and it tracks the better protocol within 3% under the large leader delay (344 vs 334 ms) and within 1% under the network-wide conditions: 684 vs 685 ms under the partial random delay, 1122 vs 1121 ms under the full one, and 1022 vs 1017 ms under jitter."

The crash dips of panel (iii) cost 7% and 17% for the two pairs, against 2% at n = 50.

"Mysticeti survives jitter at a 3.1 s plateau rather than 5.7 s, and BlueBottle-PS ties its asynchronous variant under the network-wide conditions instead of falling behind it (685, 806 and 915 ms) while its period wanders across the candidates, which costs nothing."

Claims C1 to C4 hold at n = 50 as well.

- Table Summary:

Table 3. Discharge table mapping interface clauses to its discharging lemma or base-paper result. (Details clause assignments for Mysticeti and BlueBottle variants.)

- Appendix Details:

"Appendix B states every result of Section 5 and proves it. Each proof names the clause of Section 2 at the step that relies on it; every argument that needs neither clause A4 nor clause A5 holds for any pair of rules of the wave family."

"The Lean development spans about 21k lines, partitioned so that a human reader audits only definitions and theorem statements, about 7k lines. The proofs, about 10k lines, are AI-generated, checked by the kernel and need no reading."

The model has no execution; five of the paper’s assumptions enter as hypotheses rather than as derived facts.

"New for this paper. No base formalization states either of the two new hypotheses: the coin’s unpredictability (clause A5) and the per-round floor of committed candidates that clause A5 supplies is a hypothesis the model does not derive."

- Specific Observations:

"Observation 1: Minimum-Quorum References. A validator waits up to T for the leader block before proposing, and the timer is armed at its own proposal. With every link delayed by D > T, the timer fires before any block of the round arrives, so the validator proposes the instant its threshold clock reaches n − f blocks."

Observation 2: Layers of Evidence. The second layer turns the rule into an all-or-nothing filter: the slot commits directly only if every vote-round block voted.

The counting lemma underpinning the asynchronous bound requires wa ≥ 4: at wa = 4 it guarantees only one directly committed candidate per populated round, so p ≥ 1/n.

"The counting fraction p of clause A5 is bounded per rule. For Mahi-Mahi at wa = 5 the counting lemma (Lemmas 12 and 13 of that paper) gives n − f directly committable blocks per populated round at n = 3f + 1; counting only the n − f − b honest ones among them bounds p below by (n − f − b)/n, at least 1/3 at n ≥ 3f + 1; at wa = 4 it promises only one committed candidate per populated round (its Lemma 15), so p ≥ 1/n. BlueBottle’s asynchronous variant counts at wa = 3 (Lemmas 28 to 31 of that paper): on a quorum of honest validators populating the two rounds above a round, at least n − 3f honest blocks of that round are directly committed, so p ≥ (n − 3f)/n."

The tail asks of the interval only maxPeriod ≤ interval and wa ≤ interval: at the implementation’s interval = 128 and maxPeriod = 64, a block opens every third interval at wa = 5 and every second at wa = 4.

- Conclusion:

Steelhead implements atomic broadcast (Definition 1) for any pair of rules satisfying clauses A1 to A5.

"The two readings overlap and need not agree: a round divisible by the period is a slot of both. Where either reading decides directly, both decide alike (Theorem 1); only their indirect verdicts can differ, because each searches anchors among its own slots."

A period change never has to pause the protocol to finish pending slots: the anchors of the next period decide the slots left ⊥ under the previous one.

The liveness argument of Theorem 3 uses only the control reading and the failover; at k = 1 a run of wa direct commits decides every slot below it, so every slot is eventually decided.

Under asynchrony nothing they commit is output anyway (Section 3.3), since the anchor search of an ⊥ synchronous slot below stops at the next ⊥ one.

Every output block is a block of the DAG and so carries its author’s signature (Section 2).

This summary is long and detailed, quoting relevant parts from the paper.


Steelhead: Interleaving Partially Synchronous and Asynchronous Commit Rules on a Shared DAG. George Danezis et al., arXiv:2609.30163v1 [cs.DC] 24 Sep 2026.

Steelhead: Interleaving Partially Synchronous and Asynchronous Commit Rules on a Shared DAG. George Danezis et al., arXiv:2609.30163v1 [cs.DC] 24 Sep 2026.

Improvements for AI systems

Here are the specific improvements to AI systems that can be derived from the Steelhead protocol, categorized by capability:


  1. Adaptive Consensus for Heterogeneous Network Conditions: The system gains the ability to maintain high performance across vastly different network topologies and failure modes (partial synchrony vs. asynchrony).

  2. Zero-Message Mode Switching for Protocol Selection: The AI system can dynamically switch its underlying consensus mechanism (from fast, synchronous-like operation to resilient, asynchronous operation) based purely on the observed state of the shared Directed Acyclic Graph (DAG), without needing external timers or complex voting rounds.

  3. Automated Liveness Recovery Under Asynchrony: The system is guaranteed to remain live even under severe network asynchrony (where message delays are unpredictable). It achieves this by falling back to a known, slower asynchronous rule when the fast rule stalls, ensuring no permanent deadlock occurs due to leader uncertainty.

  4. Optimized Latency Selection via Counterfactual Replay: The system can re-evaluate the optimal commit period (the dial) at runtime using only the history of committed blocks. It compares theoretical performance metrics (expected message delays under various candidate periods) against current network evidence, selecting the period that minimizes expected latency for the current conditions.

  5. Self-Tuning Period Adaptation: The protocol automatically adjusts its operational period based on whether it is currently operating in a synchronous regime (favoring smaller periods for speed) or an asynchronous regime (favoring larger periods to maintain liveness). This adaptation happens deterministically based on the committed DAG, eliminating the need for external mode votes or complex timeout counting.

  6. Robustness Against Leader Failures and Stalls: The system can gracefully handle targeted leader attacks or network stalls by utilizing an anchor search mechanism that skips undecided slots and relies on previously committed blocks above them to resolve uncertainty, preventing output from stalling entirely.

  7. Deterministic Period Agreement Across Distributed Nodes: Even when operating under highly variable conditions (like high jitter or random delays), all honest nodes in the network will compute the exact same period for every interval, ensuring a consistent view of the protocol's state and predictable behavior across the distributed system.

In summary, an AI system implementing Steelhead can function as a self-aware consensus engine capable of achieving high throughput in controlled environments while maintaining guaranteed liveness and performance under unpredictable, adversarial network conditions without requiring complex external coordination or reliance on precise clock synchronization.

Abstract

Dual-mode consensus protocols are fast when the network is partially synchronous and remain live under asynchrony. We introduce Steelhead, a dual-mode mechanism that composes a partially synchronous and an asynchronous commit rule over one DAG: every k-th round is decided by the asynchronous rule, whose leader a common coin reveals after the votes, and all other rounds by the partially synchronous rule. Every interval, validators replay the committed DAG under each candidate period, adopt the one with the fewest expected message delays, and fall back to k = 1 when the output stalls; the asynchronous rule applied to the coin rounds alone keeps the protocol live. Steelhead sends no message beyond the DAG's blocks, not even to agree on the period, and opens a coin only on the rounds that need a hidden leader. It is generic over pairs of DAG commit rules that share a committee; we instantiate it with Mysticeti and Mahi-Mahi at n >= 3f+1 and with the two variants of BlueBottle at n >= 5f+1. We prove it safe and live, and provide mechanized proofs in Lean 4. In simulation, Steelhead matches the partially synchronous protocol in a healthy network, stays close to the asynchronous one when network conditions stall the partially synchronous one, and adapts quickly in both directions.

Sources

Related papers