SLED: Scalable Location Encoding via Distillation
Kevin Lane, Zhongying Wang, Esther Rolf, Morteza Karimzadeh
cs.CV, cs.AI
Submitted: 2026-08-06
Updated: 2026-08-10
Code: https://github.com/geohai/sled
License: http://creativecommons.org/licenses/by/4.0/
The gist: The plethora of readily available geospatial data offers exciting opportunities to learn high quality representations of the planet, but the sheer size of the Earth Observations (EO), differing
Terminology
Abstract
The plethora of readily available geospatial data offers exciting opportunities to learn high quality representations of the planet, but the sheer size of the Earth Observations (EO), differing modalities, and different sensor types pose significant challenges in doing so. Location encoders have emerged as an efficient way of compressing EOs into location-specific embeddings. However, current state-of-the-art location encoders rely on computationally expensive CLIP-style frameworks that require large batch sizes in the 16K--32K range, suffer from false negative samples, and scale poorly with additional modalities. We introduce the Scalable Location Encoder via Distillation (SLED), a distillation-based location encoder that uses geospatial location as a binding modality to pretrain location encoders with any modality of geospatial data. The resulting location encoder framework is lightweight, modular, and can flexibly incorporate multiple modes, while eliminating the need for spatiotemporal coregistration of samples. SLED is performant with batch sizes as small as 128, enabling pretraining at a fraction of the runtime and compute costs of current state-of-the-art models. We demonstrate our approach by pretraining unimodal and multimodal SLED models on Sentinel-1, Sentinel-2, and Landsat imagery. We show that both unimodal and multimodal SLED models keep pace with or outperform existing approaches on a diverse set of 19 human-centric benchmark tasks and explore the benefits of using additional modes in pretraining.
Sources
- UNIGEOCLIP: Unified Geospatial Contrastive Learning
- AlphaEarth Foundations: An embedding field model for accurate and efficient global mapping from sparse label data
- TerraFM: A Scalable Foundation Model for Unified Multisensor Earth Observation
- Climplicit: Climatic Implicit Embeddings for Global Ecological Tasks
- OlmoEarth: Stable Latent Image Modeling for Multimodal Earth Observation
- Matryoshka Representation Learning
- GAIR: Location-Aware Self-Supervised Contrastive Pre-Training with Geo-Aligned Implicit Representations
- Better Together: Evaluating the Complementarity of Earth Embedding Models
- Neural Plasticity-Inspired Multimodal Foundation Model for Earth Observation
Related papers
- Loss Knows Best: Detecting Annotation Errors in Videos via Loss Trajectories
- AnchorWeave: World-Consistent Video Generation with Retrieved Local Spatial Memories
- Benchmarking the Robustness of Foundation Models for Mammography under Domain Shift
- MambaX-Net: Dual-Input Mamba-Enhanced Cross-Attention Network for Longitudinal MRI Segmentation
- TeleOCR: Navigating Document Parsing Across Digital and Camera-Captured Documents
- A Survey on Efficient Vision-Language-Action Models