Skip to main content
Leveraging Ocean Data and Models: Building the Data Infrastructure, Workforce, and Modeling Capacity to Synthesize Earth’s Ocean Observations

Leveraging Ocean Data and Models: Building the Data Infrastructure, Workforce, and Modeling Capacity to Synthesize Earth’s Ocean Observations

This is a Preprint and has not been peer reviewed. This is version 1 of this Preprint.

Add a Comment

You must log in to post a comment.


Comments

There are no comments or no comments have been made public for this article.

Downloads

Download Preprint

Authors

Susanne Elizabeth Craig , Frank E. Muller-Karger, Leonid Shumilo, Maria Tzortziou

Abstract

NASA’s aquatic and marine research enterprise within the overarching NASA Biosphere Program is entering an era of unprecedented data abundance. MODIS and the Plankton, Aerosol, Cloud, ocean Ecosystem (PACE) mission already generate ocean color data at rates that strain legacy processing and archiving workflows, and forthcoming missions will add further volume and complexity. These satellite streams are joined by rapidly growing in situ networks (e.g., BGC-Argo), suborbital campaigns, genomic surveys, citizen science, and increasingly high-resolution numerical models. The result is a Big Data regime in which the volume, variety, and acquisition rate of ocean relevant Earth system data exceed the capacity of legacy processing, storage, and synthesis approaches. An important goal is to fully integrate observing workflows, understanding, and models across the terrestrial, aquatic, and atmosphere domains to fully enable understanding of life and the realization of improved information to manage uses of ecosystem services. This white paper addresses Leveraging Ocean Data and Models, a cross-cutting Grand Challenge that underlies and enables progress on the four other foundational grand challenges under the NASA Biosphere, Atmosphere, Cryosphere, and Hydrosphere framework - Global Biosphere, Carbon and the Elements of Life, Interface Habitats, and Transient Events - each of which depends on the ability to acquire, harmonize, and synthesize disparate satellite, suborbital, in situ, and model data streams. We recommend a decadal strategy built on two complementary pillars: Access and Utility of Ocean Observational Data, covering tiered cloud-based facilities, workforce training, international data sharing, and machine learning; and Numerical Models and Data Assimilation, covering mechanistic and statistical models, coupled physical-biogeochemical assimilation, hybrid approaches, mission emulators, and a NASA Digital Twin Ocean. Without concerted, proactive investment in this cross-cutting infrastructure, the scientific and societal returns on the coming decade’s expanded observing system will be substantially diminished.

DOI

https://doi.org/10.31223/X5BV4N

Subjects

Artificial Intelligence and Robotics, Atmospheric Sciences, Biogeochemistry, Climate, Computer Sciences, Databases and Information Systems, Earth Sciences, Environmental Sciences, Fresh Water Studies, Geology, Hydrology, Meteorology, Numerical Analysis and Scientific Computing, Oceanography, Other Earth Sciences, Other Oceanography and Atmospheric Sciences and Meteorology

Keywords

ocean data infrastructure, cloud computing, machine learning, data assimilation, digital twin ocean, biogeochemical models, workforce development, ESAS 2028 decadal survey

Dates

Published: 2026-09-25 22:18

Last Updated: 2026-09-25 22:18

License

CC BY Attribution 4.0 International

Additional Metadata

Conflict of interest statement:
None

Data Availability:
No original data was used to author this paper. It is essentially a community consensus document offering guidance on the NASEM Earth Science 2028 Decadal Survey.

Metrics

Views: 36

Downloads: 2