Walrus: An Efficient Decentralized Storage Network
summary
The gist
The paper presents Walrus, a decentralized blob storage system that addresses the fundamental trade-off between replication overhead, recovery efficiency, and security guarantees in decentralized
This episode discusses
- Walrus: An Efficient Decentralized Storage Network · Paper Radio
- IPFS - Content Addressed, Versioned, P2P File System
The paper
Walrus: An Efficient Decentralized Storage Network · Read on arXiv
George Danezis, Giacomo Giuliari, Lefteris Kokoris Kogias, Markus Legner, Jean-Pierre Smith, Alberto Sonnino, Karl Wüst
Mysten Labs · University College London
Decentralized storage faces a fundamental trade-off between replication overhead, recovery efficiency, and security guarantees. Current approaches either rely on full replication, incurring substantial storage costs, or employ erasure-coding schemes that struggle with efficient recovery, especially under high churn. We present Walrus, a decentralized blob storage system that addresses these limitations through multiple technical innovations. At the core of Walrus is Red Stuff, a two-dimensional erasure-coding protocol that achieves high security with only a 4.5x replication factor, while providing self-healing of lost data. This means that recovery is done without centralized coordination and requires bandwidth proportional to the amount of lost data. However, Red Stuff on its own is not sufficient for Walrus, as it is designed with a static set of participants in mind. To further support decentralization, we also introduce a multi-stage epoch-change protocol that efficiently handles storage node churn while maintaining uninterrupted availability during committee transitions. Our system incorporates authenticated data structures to defend against malicious clients and ensure data consistency throughout storage and retrieval. Walrus has been deployed in production since March 2025 and has secured 686 TB of data by July 2026. We conduct an experimental evaluation of the deployed system and demonstrate that Walrus achieves practical performance at scale and outperforms the Arweave decentralized storage system.
Transcript
Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.
Tom: Next we'll be talking about the paper "Walrus: An Efficient Decentralized Storage Network".
Jane: The paper was written by George Danezis, Giacomo Giuliari, Lefteris Kokoris Kogias, Markus Legner, Jean-Pierre Smith et al. from Mysten Labs and University College London.
Tom: Stay tuned as we take you through the paper and discuss its implications.
Title: Tom: Welcome back to the channel, everyone. Today we're digging into a fresh arXiv paper called "Walrus: An Efficient Decentralized Storage Network." Jane, I have to say, the title alone got me curious — a storage network named after a sea mammal?
Jane: Ha, exactly. And the authors are a who's who from Mysten Labs and UCL — George Danezis, Giacomo Giuliari, Lefteris Kokoris Kogias, Markus Legner, Jean-Pierre Smith, Alberto Sonnino, Karl Wüst. These are folks who've been deep in blockchain and distributed systems for years.
Tom: So when I first saw "decentralized storage," my brain went straight to Filecoin and Arweave. But the abstract makes it clear Walrus is trying to solve a different piece of the puzzle. Jane, can you break down what makes this approach special for someone who's not a systems engineer?
Jane: Sure. Think of it like this: blockchains are great at agreeing on state, but they're terrible at storing big files — every validator has to hold a full copy, so you get a hundred to a thousand times replication. That's wasteful if all you want to do is store a video or a dataset and retrieve it later, not compute on it.
Tom: Right, so Walrus is specifically for blobs — binary large objects. And the big claim is they've built a system that's live on mainnet since March two thousand twenty-five storing hundreds of terabytes. That's not a toy.
Jane: Exactly. And the key innovation is something they call Red Stuff — a two-dimensional erasure coding scheme. I'll get into the details in a bit, but the headline is they achieve high durability with only a four point five times storage overhead, and they can recover lost data without downloading the whole file.
Tom: That recovery part is huge. In older erasure-coded systems, if a node goes down, you basically have to reconstruct the entire file to fix it. Walrus claims they can heal a single shard by pulling just a fraction of the data. That's the kind of thing that makes a system actually work in the real world.
Jane: And it's not just theory — they've deployed it. We're going to talk about the encoding scheme, the epoch changes, and the production numbers. But first, I want to flag something in the intro that I found really interesting: they explicitly compare themselves to Filecoin and Arweave, and they point out the trade-offs. Replication gives you easy recovery but costs twenty-five times overhead for twelve nines of durability. Classic erasure coding cuts that to three times but makes recovery painful.
Tom: So Walrus sits in the middle — four point five times overhead, but with recovery that's proportional to what's lost, not the whole blob. That's the sweet spot they're aiming for.
Jane: Exactly. And that's what we'll dig into next — how Red Stuff actually achieves that with a two-dimensional encoding. Stick around.
Summary: Tom: So we're back with "Walrus: An Efficient Decentralized Storage Network." Jane, last segment we teased the two-dimensional encoding. Can you walk us through how Red Stuff actually works, in plain terms?
Jane: Happy to. Imagine you have a file, and you split it into a grid — let's say rows and columns. Walrus splits the blob into a matrix of small pieces, then it erasure-encodes each column, and then it erasure-encodes each row. Each storage node gets one row and one column from that extended grid.
Tom: So each node holds a primary sliver — that's the row — and a secondary sliver — that's the column. And the magic is that if a node loses its data, it can ask other nodes for just the intersections — the symbols where their rows and columns cross.
Jane: Precisely. A node recovering its secondary sliver only needs symbols from f plus one other nodes, and each symbol is tiny — the size of the blob divided by n squared. So the total bandwidth for recovery is proportional to the blob size divided by n, not the whole blob. That's the self-healing property.
Tom: And that's a game-changer compared to the old approach where recovering one lost shard meant downloading the entire file and re-encoding it. The paper actually walks through two strawman designs first — full replication and classic erasure coding — to show why neither works well for a permissionless system with churn.
Jane: Right. Full replication is simple but costs twenty-five times storage for high durability. Classic erasure coding is efficient on storage but recovery is brutal — O of the blob size per failed node. Red Stuff splits the difference.
Tom: Now, there's a subtlety here. The paper defines a new problem called Asynchronous Complete Data Storage, or ACDS. It's not just about storing data — it's about guaranteeing that if a writer successfully writes a blob, every honest node eventually holds a piece of it, and that readers agree on what they read.
Jane: And that's where the blockchain comes in. Walrus uses Sui as a control plane — for registering blobs, publishing availability certificates, and coordinating epoch changes. The data itself flows directly between clients and storage nodes, not through the blockchain.
Tom: So the blockchain handles the metadata and the proofs, but the heavy lifting — the actual bytes — goes peer to peer. That's how they get the throughput numbers we'll talk about later.
Jane: Exactly. And there's a really clever bit about handling malicious writers. If a writer submits garbage that doesn't decode properly, the storage nodes can produce an inconsistency proof — a set of symbols and openings that prove the encoding is broken. Once enough nodes verify that, they all agree to reject the blob.
Tom: So it's not just about honest failures — it's about Byzantine behavior on both sides. Writers can be malicious, storage nodes can be malicious, and the system still maintains consistency.
Jane: Right. And that's what makes the ACDS definition interesting — it's stronger than what a lot of prior work offers. We'll talk about how that plays out in the epoch change protocol next.
Improvements: Tom: Welcome back. We're still on "Walrus: An Efficient Decentralized Storage Network." Jane, we've covered the encoding and the write/read flow. But the part that really impressed me is the epoch change — that's where most decentralized storage systems fall apart.
Jane: Oh, absolutely. The problem is simple: storage nodes join and leave. When a new committee takes over, the old nodes need to hand off their data. But if you keep writing new blobs to the old committee while they're trying to transfer petabytes to the new committee, you get a race — they either stop accepting writes or never finish the handover.
Tom: And the paper's solution is to decouple reads and writes during the transition. Writes go to the new committee immediately, but reads still go to the old committee until the new nodes have actually received their shards.
Jane: Right. And each blob carries the epoch it was written in, so clients know which committee to query. The new committee signals readiness once nodes holding two-thirds plus one of the shards have bootstrapped. Only then do reads switch over.
Tom: That's a really pragmatic design. But here's the thing — this only works efficiently because of Red Stuff. If a departing node is offline, the incoming node can't just copy the data. It has to recover it from the rest of the committee. And with Red Stuff, that recovery is cheap — proportional to the lost data, not the whole blob.
Jane: Exactly. The paper makes that point explicitly: without Red Stuff, a single faulty node would require bandwidth equal to the size of the file to be transferred across the network. That's why no prior decentralized system has managed a smooth epoch change under churn.
Tom: And they have real data to back this up. On mainnet, they've had epoch changes where shards were reassigned, and they measured transfer rates of over a gigabit per second. One case at epoch eleven moved twelve shards — about two point five terabytes total — in roughly four hours.
Jane: And they also had cases where nodes went offline entirely, forcing recovery. At epoch seventeen a node holding four shards was removed, and the new owners recovered each shard in about seventeen to twenty-one hours. The system stayed available the whole time.
Tom: That's the key claim — no downtime during reconfiguration. And the numbers suggest they've actually achieved it in production, not just in a lab.
Jane: Right. And that's a big deal for the practical adoption of decentralized storage. If you can't guarantee availability during churn, you can't run a real service on it. Walrus seems to have cracked that.
Tom: So the improvements here aren't just theoretical — they're measured on a live network with hundreds of terabytes. That's what separates this paper from a lot of academic work.
Jane: Definitely. And next we'll wrap up with the bigger picture — what this means for the future of decentralized storage.
Conclusion: Tom: And we're at the finish line for "Walrus: An Efficient Decentralized Storage Network." Jane, give us the final summary — what did we learn?
Jane: So Walrus is a production-grade decentralized blob storage system that uses a two-dimensional erasure coding scheme called Red Stuff. It achieves a four point five times storage overhead, which is far better than replication, and it can recover lost shards with bandwidth proportional to what's lost — not the whole blob. That self-healing property is what makes epoch changes and node churn manageable.
Tom: And it's not just a design — it's live. Since March two thousand twenty-five the mainnet has stored over six hundred eighty-six terabytes of data from millions of blobs. They measured read throughput of around four hundred megabytes per second for a single client, and writes at over sixty megabytes per second. Compare that to Arweave's roughly seven megabytes per second for writes.
Jane: And the latency numbers are striking too. Walrus reads for a one hundred thirty-five-megabyte blob take under five seconds. Writes take under twenty seconds. Arweave takes over half an hour just to reach finality. That's a difference of orders of magnitude.
Tom: The paper also formalizes a new problem — Asynchronous Complete Data Storage — which is a stronger guarantee than what prior work offered. And they prove their protocol satisfies all the properties: write completeness, read consistency, and validity.
Jane: Right. And the implications are significant. If decentralized storage can match centralized services on latency and throughput while offering censorship resistance and durability, it becomes a real alternative for things like media hosting, data archiving, and even AI training datasets.
Tom: There's also the cultural angle — this could enable communities to preserve their own history without relying on big tech platforms. The paper mentions inscriptions on Bitcoin as an example of people wanting to store data on decentralized networks. Walrus makes that practical.
Jane: Exactly. And the fact that it's open source and already deployed means the research is having real-world impact, not just sitting in a PDF.
Tom: So, Lu, Meng, Lalam — any final thoughts before we move on?
Lu: I think the most exciting part is that Walrus shows decentralized storage can be competitive on performance, not just on ideals. That changes the calculus for a lot of applications.
Meng: And from an engineering standpoint, the epoch change protocol is the hardest part to get right, and they've proven it works under real churn. That's a huge validation.
Lalam: The cultural impact is that communities can now own their data infrastructure without sacrificing usability. That's a meaningful step toward digital sovereignty.
Tom: Well said, everyone. That's a wrap on "Walrus: An Efficient Decentralized Storage Network." Great paper, great discussion. Next up, we've got something on the horizon that I think you'll all enjoy. Thanks for listening, and we'll see you in the next episode.
More episodes
- 2610.10768-Strategic Investment Decision Making for Value Creation in Energy Transition: A Reinforcement Learning Approach
- 2610.10858-RFChipAgent: Multi-Agentic AI Flow for Analog/RF Chip Design
- 2610.10613-Temporal transformer CAN encoder with federated lightweight heads for anomaly detection
- 2610.10616-When Routing Reveals Membership: Privacy Leakage from MoE Router Telemetry
- 2610.10655-Nullify: Null-Space Activation Steering for Training-Free LLM Unlearning
- 2610.11031-Language Modeling is Monotone Compression
- 2610.01253-Context-Aware Error Mitigation Orchestration for Hybrid Quantum Reinforcement Learning on NISQ Systems
- 2604.24201-CMGL: Confidence-guided Multi-omics Graph Learning for Cancer Subtype Classification
- 2609.34069-Towards Certificate-Driven Software Porting: A Self-Improving Agentic Harness for Scientific Program Optimization
- 2312.01221-Enabling Quantum Natural Language Processing for Hindi Language