Skip to main content
Masked Autoencoding (MAE) Outperforms Joint-Embedding Prediction (JEPA) for Frozen-Probe Very-High-Resolution Landslide Segmentation

Masked Autoencoding (MAE) Outperforms Joint-Embedding Prediction (JEPA) for Frozen-Probe Very-High-Resolution Landslide Segmentation

This is a Preprint and has not been peer reviewed. This is version 1 of this Preprint.

Add a Comment

You must log in to post a comment.


Comments

There are no comments or no comments have been made public for this article.

Downloads

Download Preprint

Authors

Sansar Raj Meena , Xiaochuan Tang, Filippo Catani 

Abstract

Mapping landslides from very-high-resolution (VHR) imagery is central to regional hazard assessment, yet supervised segmentation is constrained by the cost of expert pixel-level annotation. Self-supervised learning can exploit large unlabeled aerial archives, but which pretext objective best captures landslide morphology is unknown. Using a regional inventory from the 2023 Emilia-Romagna rainfall event, we compare two self-supervised objectives—masked autoencoding (MAE) and joint-embedding prediction (JEPA)—trained from scratch on 4.85 million unlabeled 0.2 m aerial patches and evaluated under an identical frozen-encoder protocol across label fractions from 1% to 100%. Although joint-embedding prediction has been reported to surpass masked reconstruction on ImageNet linear probing, low-label classification, and low-level tasks such as depth and counting, we find that MAE outperforms JEPA at every label fraction under this dense-segmentation protocol, and multiseed repeats confirm that the gap exceeds decoder-probe seed variance by roughly fifteenfold at the smallest label fraction and sixfold at full labels. In-domain masked pretraining reaches 3.16 times the landslide-class IoU of an architecture-matched randomly initialized frozen encoder, and controls show that the advantage cannot be explained by decoder training alone. Under a matched frozen-encoder protocol, MAE is comparable to a strong ADE20k-supervised SegFormer-B2 encoder at the smallest label fraction and clearly surpasses it as labels increase, as the supervised encoder's frozen features saturate; an end-to-end fine-tuned SegFormer-B2, reported as a strong externally pretrained reference, remains higher. These results situate masked reconstruction as a representation that better encodes the fine morphology relevant to label-efficient VHR landslide mapping.

DOI

https://doi.org/10.31223/X56N4B

Subjects

Physical Sciences and Mathematics

Keywords

self-supervised learning, masked autoencoder, joint-embedding predictive architecture, landslide mapping, very-high-resolution remote sensing, semantic segmentation, Emilia-Romagna

Dates

Published: 2026-08-15 19:39

Last Updated: 2026-08-15 19:39

License

CC BY Attribution 4.0 International

Additional Metadata

Conflict of interest statement:
None

Data Availability:
The supervised benchmark is derived from the publicly available 2023 Emilia-Romagna landslide inventory (Berti et al., 2025, ESSD, doi:10.5194/essd-17-1055-2025). The processed image–mask patches, the unlabeled pretraining corpus, model checkpoints, and train/validation split files will be deposited in a public repository with a persistent identifier upon publication.

Metrics

Views: 47

Downloads: 1