Walrus: An Efficient Decentralized Storage Network
Listen
Radio episode about this paper
Transcript
Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.
Tom: Next we'll be talking about the paper "Walrus: An Efficient Decentralized Storage Network".
Jane: The paper was written by George Danezis, Giacomo Giuliari, Lefteris Kokoris Kogias, Markus Legner, Jean-Pierre Smith et al. from Mysten Labs and University College London.
Tom: Stay tuned as we take you through the paper and discuss its implications.
Title: Tom: Welcome back to the channel, everyone. Today we're digging into a fresh arXiv paper called "Walrus: An Efficient Decentralized Storage Network." Jane, I have to say, the title alone got me curious — a storage network named after a sea mammal?
Jane: Ha, exactly. And the authors are a who's who from Mysten Labs and UCL — George Danezis, Giacomo Giuliari, Lefteris Kokoris Kogias, Markus Legner, Jean-Pierre Smith, Alberto Sonnino, Karl Wüst. These are folks who've been deep in blockchain and distributed systems for years.
Tom: So when I first saw "decentralized storage," my brain went straight to Filecoin and Arweave. But the abstract makes it clear Walrus is trying to solve a different piece of the puzzle. Jane, can you break down what makes this approach special for someone who's not a systems engineer?
Jane: Sure. Think of it like this: blockchains are great at agreeing on state, but they're terrible at storing big files — every validator has to hold a full copy, so you get a hundred to a thousand times replication. That's wasteful if all you want to do is store a video or a dataset and retrieve it later, not compute on it.
Tom: Right, so Walrus is specifically for blobs — binary large objects. And the big claim is they've built a system that's live on mainnet since March two thousand twenty-five storing hundreds of terabytes. That's not a toy.
Jane: Exactly. And the key innovation is something they call Red Stuff — a two-dimensional erasure coding scheme. I'll get into the details in a bit, but the headline is they achieve high durability with only a four point five times storage overhead, and they can recover lost data without downloading the whole file.
Tom: That recovery part is huge. In older erasure-coded systems, if a node goes down, you basically have to reconstruct the entire file to fix it. Walrus claims they can heal a single shard by pulling just a fraction of the data. That's the kind of thing that makes a system actually work in the real world.
Jane: And it's not just theory — they've deployed it. We're going to talk about the encoding scheme, the epoch changes, and the production numbers. But first, I want to flag something in the intro that I found really interesting: they explicitly compare themselves to Filecoin and Arweave, and they point out the trade-offs. Replication gives you easy recovery but costs twenty-five times overhead for twelve nines of durability. Classic erasure coding cuts that to three times but makes recovery painful.
Tom: So Walrus sits in the middle — four point five times overhead, but with recovery that's proportional to what's lost, not the whole blob. That's the sweet spot they're aiming for.
Jane: Exactly. And that's what we'll dig into next — how Red Stuff actually achieves that with a two-dimensional encoding. Stick around.
Summary: Tom: So we're back with "Walrus: An Efficient Decentralized Storage Network." Jane, last segment we teased the two-dimensional encoding. Can you walk us through how Red Stuff actually works, in plain terms?
Jane: Happy to. Imagine you have a file, and you split it into a grid — let's say rows and columns. Walrus splits the blob into a matrix of small pieces, then it erasure-encodes each column, and then it erasure-encodes each row. Each storage node gets one row and one column from that extended grid.
Tom: So each node holds a primary sliver — that's the row — and a secondary sliver — that's the column. And the magic is that if a node loses its data, it can ask other nodes for just the intersections — the symbols where their rows and columns cross.
Jane: Precisely. A node recovering its secondary sliver only needs symbols from f plus one other nodes, and each symbol is tiny — the size of the blob divided by n squared. So the total bandwidth for recovery is proportional to the blob size divided by n, not the whole blob. That's the self-healing property.
Tom: And that's a game-changer compared to the old approach where recovering one lost shard meant downloading the entire file and re-encoding it. The paper actually walks through two strawman designs first — full replication and classic erasure coding — to show why neither works well for a permissionless system with churn.
Jane: Right. Full replication is simple but costs twenty-five times storage for high durability. Classic erasure coding is efficient on storage but recovery is brutal — O of the blob size per failed node. Red Stuff splits the difference.
Tom: Now, there's a subtlety here. The paper defines a new problem called Asynchronous Complete Data Storage, or ACDS. It's not just about storing data — it's about guaranteeing that if a writer successfully writes a blob, every honest node eventually holds a piece of it, and that readers agree on what they read.
Jane: And that's where the blockchain comes in. Walrus uses Sui as a control plane — for registering blobs, publishing availability certificates, and coordinating epoch changes. The data itself flows directly between clients and storage nodes, not through the blockchain.
Tom: So the blockchain handles the metadata and the proofs, but the heavy lifting — the actual bytes — goes peer to peer. That's how they get the throughput numbers we'll talk about later.
Jane: Exactly. And there's a really clever bit about handling malicious writers. If a writer submits garbage that doesn't decode properly, the storage nodes can produce an inconsistency proof — a set of symbols and openings that prove the encoding is broken. Once enough nodes verify that, they all agree to reject the blob.
Tom: So it's not just about honest failures — it's about Byzantine behavior on both sides. Writers can be malicious, storage nodes can be malicious, and the system still maintains consistency.
Jane: Right. And that's what makes the ACDS definition interesting — it's stronger than what a lot of prior work offers. We'll talk about how that plays out in the epoch change protocol next.
Improvements: Tom: Welcome back. We're still on "Walrus: An Efficient Decentralized Storage Network." Jane, we've covered the encoding and the write/read flow. But the part that really impressed me is the epoch change — that's where most decentralized storage systems fall apart.
Jane: Oh, absolutely. The problem is simple: storage nodes join and leave. When a new committee takes over, the old nodes need to hand off their data. But if you keep writing new blobs to the old committee while they're trying to transfer petabytes to the new committee, you get a race — they either stop accepting writes or never finish the handover.
Tom: And the paper's solution is to decouple reads and writes during the transition. Writes go to the new committee immediately, but reads still go to the old committee until the new nodes have actually received their shards.
Jane: Right. And each blob carries the epoch it was written in, so clients know which committee to query. The new committee signals readiness once nodes holding two-thirds plus one of the shards have bootstrapped. Only then do reads switch over.
Tom: That's a really pragmatic design. But here's the thing — this only works efficiently because of Red Stuff. If a departing node is offline, the incoming node can't just copy the data. It has to recover it from the rest of the committee. And with Red Stuff, that recovery is cheap — proportional to the lost data, not the whole blob.
Jane: Exactly. The paper makes that point explicitly: without Red Stuff, a single faulty node would require bandwidth equal to the size of the file to be transferred across the network. That's why no prior decentralized system has managed a smooth epoch change under churn.
Tom: And they have real data to back this up. On mainnet, they've had epoch changes where shards were reassigned, and they measured transfer rates of over a gigabit per second. One case at epoch eleven moved twelve shards — about two point five terabytes total — in roughly four hours.
Jane: And they also had cases where nodes went offline entirely, forcing recovery. At epoch seventeen a node holding four shards was removed, and the new owners recovered each shard in about seventeen to twenty-one hours. The system stayed available the whole time.
Tom: That's the key claim — no downtime during reconfiguration. And the numbers suggest they've actually achieved it in production, not just in a lab.
Jane: Right. And that's a big deal for the practical adoption of decentralized storage. If you can't guarantee availability during churn, you can't run a real service on it. Walrus seems to have cracked that.
Tom: So the improvements here aren't just theoretical — they're measured on a live network with hundreds of terabytes. That's what separates this paper from a lot of academic work.
Jane: Definitely. And next we'll wrap up with the bigger picture — what this means for the future of decentralized storage.
Conclusion: Tom: And we're at the finish line for "Walrus: An Efficient Decentralized Storage Network." Jane, give us the final summary — what did we learn?
Jane: So Walrus is a production-grade decentralized blob storage system that uses a two-dimensional erasure coding scheme called Red Stuff. It achieves a four point five times storage overhead, which is far better than replication, and it can recover lost shards with bandwidth proportional to what's lost — not the whole blob. That self-healing property is what makes epoch changes and node churn manageable.
Tom: And it's not just a design — it's live. Since March two thousand twenty-five the mainnet has stored over six hundred eighty-six terabytes of data from millions of blobs. They measured read throughput of around four hundred megabytes per second for a single client, and writes at over sixty megabytes per second. Compare that to Arweave's roughly seven megabytes per second for writes.
Jane: And the latency numbers are striking too. Walrus reads for a one hundred thirty-five-megabyte blob take under five seconds. Writes take under twenty seconds. Arweave takes over half an hour just to reach finality. That's a difference of orders of magnitude.
Tom: The paper also formalizes a new problem — Asynchronous Complete Data Storage — which is a stronger guarantee than what prior work offered. And they prove their protocol satisfies all the properties: write completeness, read consistency, and validity.
Jane: Right. And the implications are significant. If decentralized storage can match centralized services on latency and throughput while offering censorship resistance and durability, it becomes a real alternative for things like media hosting, data archiving, and even AI training datasets.
Tom: There's also the cultural angle — this could enable communities to preserve their own history without relying on big tech platforms. The paper mentions inscriptions on Bitcoin as an example of people wanting to store data on decentralized networks. Walrus makes that practical.
Jane: Exactly. And the fact that it's open source and already deployed means the research is having real-world impact, not just sitting in a PDF.
Tom: So, Lu, Meng, Lalam — any final thoughts before we move on?
Lu: I think the most exciting part is that Walrus shows decentralized storage can be competitive on performance, not just on ideals. That changes the calculus for a lot of applications.
Meng: And from an engineering standpoint, the epoch change protocol is the hardest part to get right, and they've proven it works under real churn. That's a huge validation.
Lalam: The cultural impact is that communities can now own their data infrastructure without sacrificing usability. That's a meaningful step toward digital sovereignty.
Tom: Well said, everyone. That's a wrap on "Walrus: An Efficient Decentralized Storage Network." Great paper, great discussion. Next up, we've got something on the horizon that I think you'll all enjoy. Thanks for listening, and we'll see you in the next episode.
George Danezis, Giacomo Giuliari, Lefteris Kokoris Kogias, Markus Legner, Jean-Pierre Smith, Alberto Sonnino, Karl Wüst
Mysten Labs · University College London
cs.DC, cs.CR
Submitted: 2026-08-10
Comments: 15 pages, 7 figures. To be published in the Proceedings of the 2026 ACM SIGSAC Conference on Computer and Communications Security (CCS '26)
Code: https://github.com/MystenLabs/walrus
License: http://creativecommons.org/licenses/by/4.0/
Importance score: 79/100
The gist: The paper presents Walrus, a decentralized blob storage system that addresses the fundamental trade-off between replication overhead, recovery efficiency, and security guarantees in decentralized
Terminology
Summary
The paper presents Walrus, a decentralized blob storage system that addresses the fundamental trade-off between replication overhead, recovery efficiency, and security guarantees in decentralized storage. The authors identify that existing approaches either rely on full replication (incurring substantial storage costs) or employ erasure-coding schemes that struggle with efficient recovery, especially under high churn.
The paper notes that blockchains support decentralized computation through State Machine Replication (SMR), but are practically limited to applications requiring little data, since SMR requires all validators to replicate data fully, resulting in a large replication factor ranging from 100× to 1000×. Dedicated decentralized storage networks emerged to store blobs more efficiently.
The authors categorize existing protocols into two main categories:
-
Replication-based systems (Filecoin, Arweave): These offer complete availability of blobs on selected storage nodes, enabling easy access and seamless migration, but reliability hinges on the robustness of selected storage nodes. Achieving
twelve nines
of durability requires storing more than 25 copies, resulting in a 25× storage overhead. These systems also face Sybil attacks. -
Reed-Solomon (RS) encoding systems: These reduce replication requirements significantly—with n nodes, 1/3 potentially malicious, and asynchronous networks, RS encoding can achieve sufficient security with just 3× storage overhead. However, when a storage node fails and needs replacement, all existing storage nodes must send their slivers to the substitute node, resulting in O(blob) data transmitted across the network. Frequent recoveries can erode storage savings, making these systems incompatible with permissionless settings.
At the core of Walrus is Red Stuff, a novel two-dimensional (2D) encoding algorithm that is self-healing. It enables recovery of lost slivers using bandwidth proportional to the amount of lost data (O(blob/n)). The paper formally defines the problem of Asynchronous Complete Data Storage (ACDS) with three properties:
-
Write Completeness: If a writer is honest, every honest node holding a commitment eventually holds a part such that the blob can be recovered.
-
Read Consistency: Two honest readers reading a successfully written blob either both succeed and return the blob or both return ⊥.
-
Validity: If an honest writer successfully writes a blob, an honest reader holding a commitment can successfully read it.
Encoding Process: Red Stuff splits the original blob into (f+1)(2f+1) symbols visualized in an [f+1, 2f+1] matrix. It then:
-
Extends each of the 2f+1 columns (of size f+1) to n symbols, assigning each row as the primary sliver of a node (Figure 2a)
-
Extends each of the f+1 rows (of size 2f+1) with repair symbols to n symbols, creating secondary slivers (Figure 2b)
For each sliver, a vector commitment is computed over all symbols, forming the sliver commitment. Commitments for every sliver form the blob metadata, and a vector commitment to the metadata forms the blob commitment.
Storage Cost: The paper states that "a naive application of our 2D encoding in which nodes store fully extended slivers would lead to a storage amplification of 9×. However, nodes only store a sufficient number of symbols (2f+1 for primary, f+1 for secondary) that allows them to locally decode the rest of the extended sliver when they need to communicate symbols that are not stored to support another node's recovery, resulting in a factor 4.5× instead."
Sliver Healing: The self-healing property allows any storage node to recover its secondary sliver by asking f+1 storage nodes for symbols in their row that should also exist in the requesting node's extended column. Once all honest nodes have secondary slivers, any node can recover its primary sliver by asking the 2f+1 honest nodes for symbols in their column. The cost per node remains O(B/n) and total cost to recover the blob is O(B).
Walrus integrates Red Stuff with a blockchain (Sui in the implementation) serving as a control plane for metadata and governance.
Write Protocol: The writer encodes a blob using Red Stuff, derives a blob ID by hashing the blob commitment with metadata, submits a transaction on the blockchain to acquire storage space and register the blob, then sends sliver pairs to storage nodes. Nodes verify slivers against commitments and respond with signed acknowledgments. The writer collects 2f+1 signatures to form an availability certificate, which is published on-chain as the Point of Availability (PoA).
Read Protocol: The reader retrieves metadata from storage nodes, requests primary slivers, verifies each against sliver commitments, and after obtaining f+1 valid slivers, reconstructs the blob. The final verification step re-encodes the reconstructed blob and recomputes metadata—if it matches, the reader outputs the blob; otherwise, it outputs ⊥.
Handling Byzantine Faults: Storage nodes can misbehave by acknowledging slivers they don't hold or serving wrong slivers—defended through commitments. Malicious clients could upload inconsistent slivers—Walrus provides a mechanism where nodes share inconsistency proofs (fraud proofs) with other nodes, who verify by trial recovery. After f+1 attestations on-chain, all nodes respond with ⊥ to requests for the inconsistent blob's slivers.
Walrus introduces a multi-stage epoch-change protocol that efficiently handles storage node churn while maintaining uninterrupted availability. The key innovation is directing writes to the new committee (e+1) the moment reconfiguration starts while still directing reads to the old committee. The metadata of every blob includes the epoch in which it was first written, allowing clients to direct reads appropriately during the handover period. When nodes collectively holding 2f+1 shards signal readiness, the reconfiguration terminates.
The paper emphasizes: The key enabler of Walrus's ability to handle this gracefully is our Red Stuff algorithm, as it allows for the bandwidth cost of the faulty case to be on the same order as that of the fault-free case.
The paper provides formal proofs that Red Stuff satisfies all ACDS properties:
-
Write Completeness (Theorem 1): An honest writer sends at least 2f+1 correctly encoded slivers; nodes recover missing slivers through the two-dimensional recovery process, eventually all honest nodes hold both primary and secondary slivers.
-
Read Consistency (Theorem 2): Since the encoding scheme is deterministic and the last step of reading re-runs the encoding, a reader that accepts the read must output the blob. If one reader detects inconsistency, all honest readers must also detect it due to the binding property of vector commitments.
-
Validity (Theorem 3): Since at least 2f+1 honest storage nodes hold their slivers, an honest reader querying all nodes will eventually obtain f+1 valid primary slivers and reconstruct the blob.
Walrus has been deployed in production since March 2025 and had secured 686 TB of data by July 2026. The mainnet comprises 95 storage nodes in 17 countries with a total system capacity of 4.00 PB (56% used). Since launch, 17.6 million unique blobs have been registered by 9,675 wallets, with a maximum of approximately 834,000 blobs certified in one day (about 10 blobs per second) and a maximum total unencoded size of 21.0 TB certified in one day (average goodput of approximately 1.95 Gbps).
Performance Results:
-
Latency: Median write latencies remained less than 20 seconds for blobs up to 135 MB; blobs up to 135 MB were read in less than 5 seconds. This compares favorably to Filecoin (1.5 hours to write and seal, 3 hours to read sealed copy) and Arweave (35–38 minutes to store to finality).
-
Throughput: A single client achieved around 400 MB/s read throughput for 135 MB blobs and around 62 MB/s write throughput, compared to Arweave's plateau at around 7 MB/s.
-
Epoch-Change Overhead: The paper documents shard transfers and recoveries on mainnet, including a case where 12 shards were reassigned with approximately 2.5 TB total transferred at aggregate rates of roughly 1.4 Gbps average and 2.5 Gbps peak, completing in about four hours.
The paper provides a comparison of decentralized storage approaches:
Approach Replication for 10−12 Write/Read Cost Single Shard Recovery Non-blocking Epoch Change
Replication [17, 27] 25× O(nblob) O(blob) No
Classic ECC [21, 26] 3× O(blob) O(blob) No
Walrus + Red Stuff 4.5× O(blob) O(blob/n) Yes
The paper discusses IPFS, Filecoin, Arweave, Storj, Celestia, and Semi-AVID. It notes that Red Stuff builds on the Twin-code framework but differs by encoding data across differently sized dimensions and integrating authenticated data structures to achieve Write Completeness and Byzantine fault tolerance. The paper also notes a significant difference from AVID protocols: AVID protocols do not provide completeness, which is critical for Walrus.
Improvements for AI systems
Based on the paper, here are the specific improvements I can make to AI systems:
Improvement: Implement a two-dimensional erasure-coding protocol (Red Stuff) for distributed AI training datasets.
What the improved system can do:
-
Recover lost data shards using bandwidth proportional to the amount of lost data (O(blob/n)) rather than the full dataset size
-
Maintain 4.5× replication factor instead of 25× while achieving
twelve nines
durability -
Self-heal automatically when storage nodes fail, without centralized coordination
-
Reduce recovery cost from O(blob) to O(blob/n) per node—critical for petabyte-scale training corpora
Abstract
Decentralized storage faces a fundamental trade-off between replication overhead, recovery efficiency, and security guarantees. Current approaches either rely on full replication, incurring substantial storage costs, or employ erasure-coding schemes that struggle with efficient recovery, especially under high churn. We present Walrus, a decentralized blob storage system that addresses these limitations through multiple technical innovations. At the core of Walrus is Red Stuff, a two-dimensional erasure-coding protocol that achieves high security with only a 4.5x replication factor, while providing self-healing of lost data. This means that recovery is done without centralized coordination and requires bandwidth proportional to the amount of lost data. However, Red Stuff on its own is not sufficient for Walrus, as it is designed with a static set of participants in mind. To further support decentralization, we also introduce a multi-stage epoch-change protocol that efficiently handles storage node churn while maintaining uninterrupted availability during committee transitions. Our system incorporates authenticated data structures to defend against malicious clients and ensure data consistency throughout storage and retrieval. Walrus has been deployed in production since March 2025 and has secured 686 TB of data by July 2026. We conduct an experimental evaluation of the deployed system and demonstrate that Walrus achieves practical performance at scale and outperforms the Arweave decentralized storage system.
Sources
Related papers
- iScheduler: Reinforcement Learning-Driven Continual Optimization for Large-Scale Resource Investment Problems
- SAMM: Sharded Automated Market Maker
- InferScale: GPU-Native KV Injection for Personalized LLM Serving
- Vigil: Accountable Liveness against Selective Silence
- Steelhead: Interleaving Partially Synchronous and Asynchronous Commit Rules on a Shared DAG
- Pushing CPU Speech Synthesis to the Wall: Extreme Inference Tuning under Serverless Architecture and Billing