Seeing SDG 6 from space: local-scale monitoring of piped water and sewage systems across Africa using satellite imagery and self-supervised learning

arXiv:2411.19093 · cs.CV, cs.CY, cs.LG · Submitted 2026-08-23 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.

Tom: Next we'll be talking about the paper "Seeing SDG 6 from space: local-scale monitoring of piped water and sewage systems across Africa using satellite imagery and self-supervised learning".

Jane: The paper was written by Othmane Echchabi, Aya Lahlou, Nizar Talty, Josh Malcolm Manto, Tongshu Zheng et al. from Mila – Quebec AI Institute and School of Computer Science, McGill University and Department of Earth and Environmental Engineering, Columbia University and Center for Learning the Earth with Artificial Intelligence and Physics (LEAP) and Division of Natural and Applied Sciences, Duke Kunshan University.

Tom: Stay tuned as we take you through the paper and discuss its implications.

Title: Tom: Welcome back, everyone. I'm Tom, and alongside me is Jane. We're looking at a paper with a title that really says it all: "Seeing SDG six from space: local-scale monitoring of piped water and sewage systems across Africa using satellite imagery and self-supervised learning."

Jane: Tom, I love this title because it's so concrete. SDG six is the UN goal for clean water and sanitation, and this paper is literally trying to see that goal from orbit. The authors are using satellites to figure out where piped water and sewage systems actually exist across Africa.

Tom: And the key word there is "actually." Right now, we track these things through household surveys, which are expensive and slow. The paper says a survey can cost around three hundred dollars per household in sub-Saharan Africa. So countries just don't do them very often.

Jane: That's the gap this paper wants to fill. Instead of knocking on doors, they're looking at satellite images. But here's the clever part, Tom — they're not just using the images directly. They're using something called self-supervised learning to teach a computer what infrastructure looks like from above.

Tom: So the computer learns patterns on its own, without someone labeling millions of images by hand. That's the "self-supervised" part. And it's a big deal because labeling satellite images of the entire African continent would take forever.

Jane: Exactly. And the authors are from Duke Kunshan University, with collaborators at McGill and Columbia. They've got a mix of environmental engineering and computer science backgrounds, which makes sense for this kind of work.

Tom: The implications here are huge for how we track progress on global development goals. If this works, we could have fresh estimates every few months instead of every few years. That's the promise we're going to dig into today.

Jane: And we'll get into how well it actually works, because there are some real numbers in this paper that are pretty impressive. Stick around.

Summary: Tom: So Jane, let's get into what this paper actually did. The team built a system that looks at Sentinel-two satellite images — those are freely available, high-resolution images from the European Space Agency — and tries to predict whether a given area has piped water and sewage systems.

Jane: And the way they trained it is fascinating. They used Afrobarometer survey data, which is this massive dataset of interviews across Africa. But instead of using the household answers, they used the enumerators' observations — the field workers who noted whether the area had these systems.

Tom: Right, so you've got about eighteen thousand survey locations across forty countries, each one paired with a satellite image. The model learns what the built environment looks like when these systems are present versus when they're not.

Jane: And the results on held-out data are strong. They got an AUROC of about ninety-one and a half percent for piped water, and over ninety-three percent for sewage. For folks who don't speak statistics, that means the model is very good at telling apart places with and without these systems.

Tom: But here's what really got me excited, Jane. When they applied this across all fifty African countries and compared their national estimates to the official WHO and UNICEF numbers, the piped water estimates matched really well — an R-squared of zero point nine two.

Jane: That's the external validation, which is the gold standard. You train on one set of data, then you check against completely independent official statistics. And it held up.

Tom: Now, the sewage numbers were a bit weaker — an R-squared of zero point seven two. But the paper explains that the comparison isn't perfect because the official sanitation numbers include things like latrines, not just sewage systems.

Jane: So it's comparing apples to a slightly different fruit. But even so, the fact that satellite imagery alone can get this close to official statistics is remarkable. It means we might have a new tool for monitoring progress on clean water and sanitation.

Tom: And that's the big picture. This could complement the traditional surveys, filling in the gaps between them. We'll talk about what that means for real people in the next segment.

Improvements: Tom: Jane, one of the things I appreciate about this paper is that they didn't just stop at "it works." They dug into what makes it work and how to make it better. And they found some really interesting things about the model architecture.

Jane: Right, they compared different self-supervised models. They trained their own DINO and DINOv2 models on over a million satellite images from Africa, and they also tested some off-the-shelf models like DINOv3 and a couple of Earth observation foundation models.

Tom: And the models they trained themselves did better. Their DINO model got that ninety-one percent AUROC for piped water, while the off-the-shelf DINOv3 got about eighty-eight percent. That's a real difference.

Jane: The paper suggests that pretraining on African satellite imagery specifically matters. The off-the-shelf models were trained on more general images, so they don't capture the specific visual patterns of African settlements as well.

Tom: But here's the part I found most interesting, Jane. They tested whether the model was just learning "city versus countryside" — because cities obviously have more infrastructure. So they compared their model against a simple urban-rural indicator.

Jane: And the satellite model crushed it. The image embeddings beat the urban-rural baseline by about fifteen percentage points of AUROC. And even when they looked at only urban areas or only rural areas separately, the model still performed well.

Tom: That's the key improvement this paper demonstrates. It's not just detecting urbanization — it's picking up on actual infrastructure-related patterns. Things like road networks, building density, settlement layout. The self-supervised learning is capturing something genuinely useful.

Jane: They also tested how well the model transfers to regions it's never seen. They held out entire UN subregions of Africa during training, then tested on them. The performance dropped to about seventy-five percent AUROC for piped water, which is lower but still well above chance.

Tom: So it can generalize to new areas, but there's a cost. That's an honest limitation, and it tells us where the model needs more work — probably more diverse training data or some kind of adaptation when moving to new regions.

Jane: And that's the kind of nuance that makes this paper solid. They're not overselling it. They're showing where it works, where it struggles, and what could improve it. Let's get into the actual numbers on the first page in a moment.

First Page: Tom: So Jane, let's zoom in on the first page of "Seeing SDG six from space." The abstract lays out the whole story, and there's one number that really jumps out at me: the scale of the problem they're trying to solve.

Jane: Absolutely. The paper opens with the fact that over two billion people lack safely managed drinking water, and over three billion lack safely managed sanitation. And sub-Saharan Africa is hit hardest — only about a third of the population there has safely managed water.

Tom: That's the motivation. But the first page also makes a really practical point about why the current monitoring system is failing. Surveys are expensive and infrequent. Some countries go a decade between comprehensive water and sanitation surveys.

Jane: And by the time the data is processed and published, it's two to four years old. So you're making decisions about where to build infrastructure based on information that's already outdated.

Tom: The paper calls this a "temporal resolution mismatch." The surveys just can't keep up with how fast access changes. And that's where the satellite approach comes in — it can be refreshed constantly.

Jane: There's also a great point about the cost. The paper mentions that conducting the surveys needed across low-income countries by two thousand thirty would cost nearly a billion dollars. That's just not going to happen. So we need a cheaper way to monitor progress.

Tom: And that's what this framework offers. The satellite data is free. The self-supervised learning means you don't need massive labeled datasets. The whole pipeline is designed to be low-cost and scalable.

Jane: The first page also positions this against previous work. There's been research on using satellite imagery to predict poverty and economic well-being in Africa, but this paper specifically targets water and sanitation infrastructure, which is more directly tied to SDG six.

Tom: And they're building on that earlier work with modern techniques. The Vision Transformers and self-supervised learning are a big step up from the older convolutional neural networks used in previous studies.

Jane: So the first page sets up the problem, explains why it matters, and hints at the solution. It's a strong opening that makes you want to keep reading — and we've been doing exactly that.

Conclusion: Tom: Well, Jane, we've covered a lot of ground on "Seeing SDG six from space." Let's wrap this up. The core idea is that satellite imagery, combined with self-supervised learning, can estimate where piped water and sewage systems exist across Africa at a very fine scale.

Jane: And the results are genuinely impressive. Ninety-one percent AUROC for piped water on held-out data, and the national estimates track official statistics really well for piped water. Even in countries with no survey data at all, the errors were around ten percent.

Tom: That's the part I keep coming back to — the countries without survey coverage. The paper showed that for nearly two hundred million people in those countries, the model's estimates were within fifteen percent of official numbers for piped water. That's a huge win for monitoring.

Jane: And the Nigeria case study showed how this can be used in practice. They mapped burden and severity across over seven hundred local government areas, identifying exactly where the largest populations live without these systems. That's actionable information for planners.

Tom: The limitations are real, though. The model measures system presence, not actual household use. And the sewage comparison to official statistics is weaker because the definitions don't perfectly align. But as a complement to surveys, this is powerful.

Jane: I think the biggest takeaway is that we now have a tool that can monitor infrastructure progress continuously, at low cost, and at a local scale. That could change how we track progress on clean water and sanitation goals.

Tom: And it could be adapted to other infrastructure too — energy, roads, maybe even schools. The framework is general. That's the exciting part for the future.

Jane: Alright, Tom, I think we've given our listeners a solid picture of this paper. Thanks for joining us, everyone. Next up, we've got another paper on using machine learning for climate modeling, so stay tuned.

Tom: See you then.

Othmane Echchabi, Aya Lahlou, Nizar Talty, Josh Malcolm Manto, Tongshu Zheng, Ka Leung Lam

Mila – Quebec AI Institute · School of Computer Science, McGill University · Department of Earth and Environmental Engineering, Columbia University · Center for Learning the Earth with Artificial Intelligence and Physics (LEAP) · Division of Natural and Applied Sciences, Duke Kunshan University

cs.CV, cs.CY, cs.LG

Submitted: 2026-08-23

Updated: 2026-08-25

Comments: Under Review

Code: https://github.com/othmaneechc/SDG6Tracker

License: http://creativecommons.org/licenses/by-nc-nd/4.0/

Importance score: 71/100

The gist: This study develops a scalable remote-sensing framework for estimating area-level presence of piped water and sewage systems at 2.56 km spatial resolution across Africa.

Terminology

Summary

This study develops a scalable remote-sensing framework for estimating area-level presence of piped water and sewage systems at 2.56 km spatial resolution across Africa. The framework integrates Sentinel-2 imagery, enumerator-observed system-presence records from Afrobarometer enumeration areas, 30 m population data, and Vision Transformer representations learned with DINO self-supervised learning. The best-performing models achieve held-out AUROC values of 91.54% for piped water and 93.24% for sewage across enumeration areas. Under leave-one-region-out cross-validation, this falls to 75.5% and 78.7% respectively, reflecting transfer difficulty to unsampled regions. Applied across 50 African countries, population-weighted piped water estimates closely track WHO/UNICEF JMP piped water access (R 2 = 0.92), while sewage estimates show meaningful agreement with the broader JMP safely managed sanitation benchmark (R 2 = 0.72). In countries without Afrobarometer survey coverage, the model achieves population-weighted mean absolute errors of 9.5% for piped water and 10.7% for sewage. A Nigeria application across 767 Local Government Areas shows how our framework’s fine-scale predictions reveal substantial subnational inequality, with the largest populations living where no piped water system is present reaching 1.187 million, and no sewage system 1.577 million. These findings show that DINO-based self-supervised learning using freely available satellite imagery can complement traditional household surveys, supporting SDG 6 monitoring, infrastructure planning, and environmental equity assessment.

Improvements for AI systems

Based on the paper, here are the specific improvements I can make to an AI system and what the improved system can do:

  • Implement DINO/DINOv2 self-supervised learning on large-scale unlabeled Sentinel-2 imagery (over 1.4 million patches) before any supervised fine-tuning

  • Use Vision Transformer (ViT-Base/ViT-Large) backbones with patch size 8

  • Train with multi-crop strategy (global + local crops) and exponential moving average teacher network

  • This eliminates the need for expensive labeled data while learning transferable built-environment features

  • Replace heavy supervised classifiers with a k-nearest-neighbor classifier using cosine similarity and softmax weighting (temperature 0.07)

  • Tune k ∈ 5, 10, 20, 50, 100, 200 per task

  • This fits essentially no parameters to labels, making the system robust to small labeled datasets

  • Generate patch-level predictions at 2.56 km resolution

  • Combine with 30m gridded population data (Meta/CIESIN) to produce population-weighted estimates

  • Aggregate to national or subnational (e.g., Local Government Area) levels for policy-relevant outputs

  • Apply isotonic regression for probability calibration (reduces expected calibration error by 50%)

  • Use bootstrap resampling to provide confidence intervals on aggregated burden estimates

  • Report both burden (expected people affected) and severity (probability of absence) metrics

  • Implement leave-one-region-out cross-validation (over UN African subregions) to quantify transferability

  • This provides honest estimates of performance in unsampled regions

  • Predict presence/absence of piped water and sewage systems at 2.56 km resolution across 50 African countries

  • Achieve AUROC of 91.5% (piped water) and 93.2% (sewage) on held-out locations

  • Maintain AUROC of 75.5% and 78.7% even when entire geographic regions are held out

  • Produce national-level estimates that track WHO/UNICEF JMP statistics with R2 = 0.92 (piped water) and R2 = 0.72 (sanitation)

  • Work in countries with no survey coverage, achieving mean absolute errors of 9.5% (piped water) and 10.7% (sewage)

  • For any country (demonstrated on Nigeria's 767 Local Government Areas), identify:

  • High-burden areas: where largest populations lack infrastructure (up to 1.19M people without piped water, 1.58M without sewage in worst LGAs)

  • High-severity areas: where absence probability is most pervasive (90th percentile absence probability of 0.95 for sewage)

  • Combine both metrics to prioritize areas needing urgent intervention

  • Refresh estimates as frequently as new satellite imagery is available (2-3 day revisit for Sentinel-2)

  • Near-zero marginal cost per refresh compared to 300/household surveys

  • Complement household surveys by filling temporal and spatial gaps

  • The self-supervised pretraining approach transfers to other geographic contexts with comparable satellite coverage

  • Adaptable to other spatially observable SDG indicators (energy infrastructure, urban expansion, land productivity)

  • Give confidence intervals on burden estimates (e.g., 95% CI for max LGA burden: 0.87–1.48 million for piped water)

  • Flag where predictions are less reliable (e.g., where visible built-environment proxies are weak)

Abstract

Access to drinking water and sanitation is essential, yet monitoring progress toward Sustainable Development Goal 6 remains constrained by costly, infrequent, and spatially uneven household surveys, particularly in data-scarce regions. We develop a scalable remote-sensing framework to estimate the area-level presence of piped water and sewage systems across Africa at 2.56 km resolution. The framework combines Sentinel-2 imagery, enumerator-observed system-presence records from Afrobarometer enumeration areas, 30 m population data, and Vision Transformer representations learned through DINO self-supervised learning. On held-out enumeration areas, the best models achieve AUROCs of 91.54% for piped water and 93.24% for sewage. Under leave-one-region-out cross-validation, performance declines to 75.5% and 78.7%, respectively, indicating challenges in transferring models to unsampled regions. Applied across 50 African countries, population-weighted estimates closely track WHO/UNICEF Joint Monitoring Programme benchmarks for piped water access (R squared = 0.92) and show meaningful agreement with safely managed sanitation for sewage (R squared = 0.72). In countries without Afrobarometer coverage, population-weighted mean absolute errors are 9.5% for piped water and 10.7% for sewage. Predictions for 767 Local Government Areas in Nigeria reveal substantial subnational inequality: in the most affected areas, as many as 1.187 million people live where no piped water system is present and 1.577 million where no sewage system is present. These findings show that self-supervised learning with freely available satellite imagery can complement household surveys and support SDG 6 monitoring, infrastructure planning, and environmental equity assessment.

Related papers