This is a Preprint and has not been peer reviewed. This is version 1 of this Preprint.
Three gaps for foundation models to address in satellite remote sensing of inland waters
Downloads
Authors
Abstract
Long-term Earth observation of inland waters depends on continuity between satellite missions, adequate combinations of spectral, spatial and temporal resolution, and representative calibration and validation datasets. Discontinuities and trade-offs in these components constrain the records available for studying inland waters and their change. We examine three gaps affecting inland water remote sensing: a mission gap between comparable 300 m ocean colour observations from MERIS and OLCI during 2012–2016, a coverage gap in spectral, spatial and temporal sampling, and a generalisation gap when retrieval algorithm accuracy declines beyond bio-optical properties and atmospheric conditions represented during development. Foundation models (FMs) could help address these gaps by learning reusable representations from satellite archives before adaptation to a particular retrieval task. We hypothesise that relationships learned across overlapping satellite missions can support reconstruction of what a sensor would have measured during the mission gap, that combining complementary sensors can extend usable observation coverage, and that pretraining on satellite archives, which cover more water bodies and a wider range of bio-optical properties and atmospheric conditions than available field datasets, can improve retrieval algorithm generalisation. Testing these hypotheses requires criteria that define progress. We quantify the three gaps, define a progress criterion for each and assess 100 recent machine learning studies against the criteria. For the mission gap, three studies address continuity across sensor eras, and none recovers a record from a withheld mission period. For the coverage gap, 23 studies combine complementary observations, but only three test them on withheld water bodies or periods. For the generalisation gap, only seven studies withhold whole water bodies or later periods from all model development. Across all three gaps, none of the six studies that use pretraining compares it with the same architecture trained without pretraining on independent water bodies. The assumption common to all three hypotheses, that archive pretraining adds information beyond established methods, therefore remains untested. Applying the criteria to existing generic Earth observation FMs, then to FMs pretrained on inland waters, would determine which gaps pretraining can reduce.
DOI
https://doi.org/10.31223/X5151T
Subjects
Artificial Intelligence and Robotics, Computer Sciences, Earth Sciences, Environmental Sciences, Fresh Water Studies, Oceanography and Atmospheric Sciences and Meteorology
Keywords
Foundation models, Self-supervised learning, Pretraining, Mission gap, Generalisation, Uncertainty quantification, Sentinel, Landsat, Inland water remote sensing, Self-supervised learning, Pretraining, Mission gap, Generalisation, Uncertainty quantification, Sentinel, Landsat
Dates
Published: 2026-10-09 10:13
Last Updated: 2026-10-09 10:13
License
CC BY Attribution 4.0 International
Additional Metadata
Conflict of interest statement:
None
Data Availability:
All data used in this study are openly available. Spectral ambiguity (Figure 2A) was calculated from the SD_hyper synthetic library (Pitarch and Brando, 2025) and the sensor spectral response functions published by ESA, NASA and USGS. Lake resolvability (Figure 2B, Section 5) was calculated from HydroLAKES version 1.0 (Messager et al., 2016; https://www.hydrosheds.org/products/hydrolakes). Potential satellite observations in Figure 3 are based on the global lake area of Verpoorter et al. (2014) and nominal mission specifications, and the field and simulated dataset sizes were taken from their data publications (Bi and Xu, 2026; Castagna et al., 2022; Drayson et al., 2022; Gege and Dekker, 2020; Hanly et al., 2025; Kravitz et al., 2021; Lavigne et al., 2022; Lehmann et al., 2023; Maciel et al., 2025; Ross et al., 2019; Zhai et al., 2024). Studies for the literature assessment were identified through OpenAlex (https://openalex.org) with the search strings in Appendix A1.2.
Metrics
Views: 21
Downloads: 1
There are no comments or no comments have been made public for this article.