This is a Preprint and has not been peer reviewed. This is version 1 of this Preprint.
Leveraging Ocean Data and Models: Building the Data Infrastructure, Workforce, and Modeling Capacity to Synthesize Earth’s Ocean Observations
Downloads
Authors
Abstract
NASA’s aquatic and marine research enterprise within the overarching NASA Biosphere Program is entering an era of unprecedented data abundance. MODIS and the Plankton, Aerosol, Cloud, ocean Ecosystem (PACE) mission already generate ocean color data at rates that strain legacy processing and archiving workflows, and forthcoming missions will add further volume and complexity. These satellite streams are joined by rapidly growing in situ networks (e.g., BGC-Argo), suborbital campaigns, genomic surveys, citizen science, and increasingly high-resolution numerical models. The result is a Big Data regime in which the volume, variety, and acquisition rate of ocean relevant Earth system data exceed the capacity of legacy processing, storage, and synthesis approaches. An important goal is to fully integrate observing workflows, understanding, and models across the terrestrial, aquatic, and atmosphere domains to fully enable understanding of life and the realization of improved information to manage uses of ecosystem services. This white paper addresses Leveraging Ocean Data and Models, a cross-cutting Grand Challenge that underlies and enables progress on the four other foundational grand challenges under the NASA Biosphere, Atmosphere, Cryosphere, and Hydrosphere framework - Global Biosphere, Carbon and the Elements of Life, Interface Habitats, and Transient Events - each of which depends on the ability to acquire, harmonize, and synthesize disparate satellite, suborbital, in situ, and model data streams. We recommend a decadal strategy built on two complementary pillars: Access and Utility of Ocean Observational Data, covering tiered cloud-based facilities, workforce training, international data sharing, and machine learning; and Numerical Models and Data Assimilation, covering mechanistic and statistical models, coupled physical-biogeochemical assimilation, hybrid approaches, mission emulators, and a NASA Digital Twin Ocean. Without concerted, proactive investment in this cross-cutting infrastructure, the scientific and societal returns on the coming decade’s expanded observing system will be substantially diminished.
DOI
https://doi.org/10.31223/X5BV4N
Subjects
Artificial Intelligence and Robotics, Atmospheric Sciences, Biogeochemistry, Climate, Computer Sciences, Databases and Information Systems, Earth Sciences, Environmental Sciences, Fresh Water Studies, Geology, Hydrology, Meteorology, Numerical Analysis and Scientific Computing, Oceanography, Other Earth Sciences, Other Oceanography and Atmospheric Sciences and Meteorology
Keywords
ocean data infrastructure, cloud computing, machine learning, data assimilation, digital twin ocean, biogeochemical models, workforce development, ESAS 2028 decadal survey
Dates
Published: 2026-09-25 22:18
Last Updated: 2026-09-25 22:18
License
CC BY Attribution 4.0 International
Additional Metadata
Conflict of interest statement:
None
Data Availability:
No original data was used to author this paper. It is essentially a community consensus document offering guidance on the NASEM Earth Science 2028 Decadal Survey.
Metrics
Views: 36
Downloads: 2
There are no comments or no comments have been made public for this article.