Skip to main content
Agentic Large Language Model Workflows for Earth Observation: Tool Orchestration, Foundation-Model Adaptation, and the Evaluation Gap

Agentic Large Language Model Workflows for Earth Observation: Tool Orchestration, Foundation-Model Adaptation, and the Evaluation Gap

This is a Preprint and has not been peer reviewed. This is version 1 of this Preprint.

Add a Comment

You must log in to post a comment.


Comments

There are no comments or no comments have been made public for this article.

Downloads

Download Preprint

Authors

salah Mohammed alheejawi , Valentas Gruzauskas

Abstract

Large language model (LLM) agents have moved rapidly from general-purpose assistants to domain systems that orchestrate specialised remote-sensing models. Within roughly two years, geospatial agents have demonstrated reliable task planning across dozens of Earth Observation (EO) tasks, benchmark environments containing hundreds of tools, and fine-tuned open-weight controllers that outperform far larger proprietary models at generating geospatial tool-use chains. This review synthesises that literature and argues that it has converged on a common architecture with a common limit. Every system surveyed orchestrates models that human researchers built, over data that human analysts prepared; none designs a model, and none constructs its own analysis-ready data. We identify three gaps that follow from this. First, a reliability wall: reported agent success falls from above 95% on single-tool tasks to below 70% beyond eight tool calls, and scaling benchmark volume tenfold does not expose or improve this, so the binding variable is composition depth rather than data quantity. Second, a preprocessing gap: benchmark datasets arrive co-registered, temporally aligned and quality-flagged, so agents are never evaluated on the multi-source data-cube construction that dominates real EO practice. Third, an evaluation gap: automated LLM-as-judge scoring is widely reported as ground truth without calibration against human experts. We survey the open geospatial foundation models that now form the practical substrate for EO modelling, argue that adaptation strategy rather than architecture is the under-searched design variable, and set out a research agenda in which domain constraint — verified templates, physically grounded data contracts, explicit compute budgets and the capacity to refuse — is treated as the mechanism by which agentic autonomy becomes verifiable rather than merely fluent.

DOI

https://doi.org/10.31223/X5TN59

Subjects

Engineering

Keywords

Earth Observation, large language model agents, remote sensing, geospatial foundation models, multi-sensor data fusion, agent evaluation, trustworthy AI

Dates

Published: 2026-09-05 23:39

Last Updated: 2026-09-05 23:39

License

CC BY Attribution 4.0 International

Additional Metadata

Conflict of interest statement:
None

Data Availability:
This is a review article; no new primary data were generated. no code or data accompany this review.

Metrics

Views: 30

Downloads: 0