This is a Preprint and has not been peer reviewed. This is version 1 of this Preprint.
Agentic Large Language Model Workflows for Earth Observation: Tool Orchestration, Foundation-Model Adaptation, and the Evaluation Gap
Downloads
Authors
Abstract
Large language model (LLM) agents have moved rapidly from general-purpose assistants to domain systems that orchestrate specialised remote-sensing models. Within roughly two years, geospatial agents have demonstrated reliable task planning across dozens of Earth Observation (EO) tasks, benchmark environments containing hundreds of tools, and fine-tuned open-weight controllers that outperform far larger proprietary models at generating geospatial tool-use chains. This review synthesises that literature and argues that it has converged on a common architecture with a common limit. Every system surveyed orchestrates models that human researchers built, over data that human analysts prepared; none designs a model, and none constructs its own analysis-ready data. We identify three gaps that follow from this. First, a reliability wall: reported agent success falls from above 95% on single-tool tasks to below 70% beyond eight tool calls, and scaling benchmark volume tenfold does not expose or improve this, so the binding variable is composition depth rather than data quantity. Second, a preprocessing gap: benchmark datasets arrive co-registered, temporally aligned and quality-flagged, so agents are never evaluated on the multi-source data-cube construction that dominates real EO practice. Third, an evaluation gap: automated LLM-as-judge scoring is widely reported as ground truth without calibration against human experts. We survey the open geospatial foundation models that now form the practical substrate for EO modelling, argue that adaptation strategy rather than architecture is the under-searched design variable, and set out a research agenda in which domain constraint — verified templates, physically grounded data contracts, explicit compute budgets and the capacity to refuse — is treated as the mechanism by which agentic autonomy becomes verifiable rather than merely fluent.
DOI
https://doi.org/10.31223/X5TN59
Subjects
Engineering
Keywords
Earth Observation, large language model agents, remote sensing, geospatial foundation models, multi-sensor data fusion, agent evaluation, trustworthy AI
Dates
Published: 2026-09-05 23:39
Last Updated: 2026-09-05 23:39
License
CC BY Attribution 4.0 International
Additional Metadata
Conflict of interest statement:
None
Data Availability:
This is a review article; no new primary data were generated. no code or data accompany this review.
Metrics
Views: 30
Downloads: 0
There are no comments or no comments have been made public for this article.