This is a Preprint and has not been peer reviewed. This is version 1 of this Preprint.
Masked Autoencoding (MAE) Outperforms Joint-Embedding Prediction (JEPA) for Frozen-Probe Very-High-Resolution Landslide Segmentation
Downloads
Authors
Abstract
Mapping landslides from very-high-resolution (VHR) imagery is central to regional hazard assessment, yet supervised segmentation is constrained by the cost of expert pixel-level annotation. Self-supervised learning can exploit large unlabeled aerial archives, but which pretext objective best captures landslide morphology is unknown. Using a regional inventory from the 2023 Emilia-Romagna rainfall event, we compare two self-supervised objectives—masked autoencoding (MAE) and joint-embedding prediction (JEPA)—trained from scratch on 4.85 million unlabeled 0.2 m aerial patches and evaluated under an identical frozen-encoder protocol across label fractions from 1% to 100%. Although joint-embedding prediction has been reported to surpass masked reconstruction on ImageNet linear probing, low-label classification, and low-level tasks such as depth and counting, we find that MAE outperforms JEPA at every label fraction under this dense-segmentation protocol, and multiseed repeats confirm that the gap exceeds decoder-probe seed variance by roughly fifteenfold at the smallest label fraction and sixfold at full labels. In-domain masked pretraining reaches 3.16 times the landslide-class IoU of an architecture-matched randomly initialized frozen encoder, and controls show that the advantage cannot be explained by decoder training alone. Under a matched frozen-encoder protocol, MAE is comparable to a strong ADE20k-supervised SegFormer-B2 encoder at the smallest label fraction and clearly surpasses it as labels increase, as the supervised encoder's frozen features saturate; an end-to-end fine-tuned SegFormer-B2, reported as a strong externally pretrained reference, remains higher. These results situate masked reconstruction as a representation that better encodes the fine morphology relevant to label-efficient VHR landslide mapping.
DOI
https://doi.org/10.31223/X56N4B
Subjects
Physical Sciences and Mathematics
Keywords
self-supervised learning, masked autoencoder, joint-embedding predictive architecture, landslide mapping, very-high-resolution remote sensing, semantic segmentation, Emilia-Romagna
Dates
Published: 2026-08-15 19:39
Last Updated: 2026-08-15 19:39
License
CC BY Attribution 4.0 International
Additional Metadata
Conflict of interest statement:
None
Data Availability:
The supervised benchmark is derived from the publicly available 2023 Emilia-Romagna landslide inventory (Berti et al., 2025, ESSD, doi:10.5194/essd-17-1055-2025). The processed image–mask patches, the unlabeled pretraining corpus, model checkpoints, and train/validation split files will be deposited in a public repository with a persistent identifier upon publication.
Metrics
Views: 47
Downloads: 1
There are no comments or no comments have been made public for this article.