Skip to main content
Evaluating forecast baselines and network representativeness for daily PM2.5 in the Kathmandu Valley, Nepal

Evaluating forecast baselines and network representativeness for daily PM2.5 in the Kathmandu Valley, Nepal

This is a Preprint and has not been peer reviewed. This is version 1 of this Preprint.

Add a Comment

You must log in to post a comment.


Comments

There are no comments or no comments have been made public for this article.

Downloads

Download Preprint

Authors

Buddhi Sagar Poudel , Nita Khatri, Namrata Mishra, Dhiraj Pradhananga

Abstract

An order-fair attribution of daily PM2.5 predictability in the Kathmandu Valley assigns 72% of achievable
error reduction to persistence, 19% to meteorology and 9% to day-of-year climatology, with the latter two
largely redundant; no fitted model, including random forests and gradient boosting tuned by nested
chronological cross-validation, outperforms an unfitted lag-1 forecast (11.22 µg m-3). We establish this
using four years (2021-2024; 5,138 station-days) of daily PM2.5, PM10 and total suspended particulates
from six Department of Environment stations, together with meteorological observations including an
hourly 10 m wind record, to test two assumptions common in machine-learning studies of this basin,
which report coefficients of determination of 0.80-0.91 pooled across seasons without naive baselines and
treat the monitoring network as a set of independent sampling points. Because a three-month season
retains a strong sub-seasonal trend, day-to-day relationships are evaluated after removing a day-of-year
climatology from both response and predictor, fitted on training data only; confidence intervals account
for temporal autocorrelation and for clustering within six stations. Cross-station correlation of daily log
PM2.5 is 0.86 (95% CI 0.79-0.91), falling to 0.61 (0.52-0.72) after deseasonalisation, with no decay
detectable over the 2-13 km network span, though the design resolves only decay larger than 0.11-0.21 r-
units. A pooled model attributes 71% of permutation importance to minimum temperature; after
deseasonalisation, held-out R2 falls from 0.64 to 0.10 and that importance disappears. Predictive skill in
this basin should be reported against climatology and persistence baselines, and pooled meteorological
importances should not be read as physical attribution.

DOI

https://doi.org/10.31223/X5FJ67

Subjects

Atmospheric Sciences, Climate, Environmental Indicators and Impact Assessment, Environmental Monitoring, Meteorology

Keywords

PM2.5, Kathmandu Valley, spatial representativeness, forecast baselines, persistence, seasonal confounding

Dates

Published: 2026-09-03 21:09

Last Updated: 2026-09-03 21:09

License

CC BY Attribution 4.0 International

Additional Metadata

Data Availability:
The merged analysis dataset, station crosswalk, and analysis and figure-generation code are archived on Zenodo (https://doi.org/10.5281/zenodo.22148941); files are available on request through the Zenodo platform.

Metrics

Views: 19

Downloads: 3