Skip to main content
Kaya-guided hybrid deep learning for leave-country-out estimation of national CO2 emissions: a 112-country benchmark with country-clustered inference, conformal uncertainty, and explainable AI

Kaya-guided hybrid deep learning for leave-country-out estimation of national CO2 emissions: a 112-country benchmark with country-clustered inference, conformal uncertainty, and explainable AI

This is a Preprint and has not been peer reviewed. This is version 1 of this Preprint.

Add a Comment

You must log in to post a comment.


Comments

There are no comments or no comments have been made public for this article.

Downloads

Download Preprint

Supplementary Files

Authors

SAMIUEL C. ATUAHENE , Bismark Boateng, Hilda Amarfio , Ekow Yaleh Quayson, Francis Boateng Agyenim

Abstract

Reliable national CO2 inventories underpin the Paris Agreement’s Enhanced Transparency 1national emissions modelling as leave-country-out (LOCO) estimation: predicting a country’s per-capita CO2 from eight-year trajectories of energy and socioeconomic covariates, with no access to its own emission history. Using Global Carbon Budget and Energy Institute data harmonized by Our World in Data, we build a 112-country panel (1990-2023; 3,024 windowed samples) and benchmark thirteen models under country-disjoint five-fold cross-validation. The proposed Kaya-guided CNN-BiLSTM-Attention network (KG-CBA) couples a log-linear backbone motivated by the Kaya identity to a convolutional-recurrent-attention branch that learns only the nonlinear residual, tuned by particle swarm optimization. Because the sample contains 27 repeated years per country, all inference is country-clustered. Kaya guidance significantly improves the unguided architecture (median APE 15.7% vs 20.8%, p = 0.002), and ridge regression significantly outperforms every tree ensemble and every unguided deep network; but the best ensemble’s apparent edge over ridge is not significant (15.1% vs 16.3%, p = 0.22), where observation-level testing would report p < 10-9. Error is highly concentrated: six structurally atypical energy systems contribute 79-84% of total squared error. Adding the low-carbon energy share on a 76-country sub-panel halves median error (7.35% vs 13.72% like-for-like), showing that energy-mix data availability, not model sophistication, is the binding constraint. Removing all CO2-derived covariates costs under one percentage point, so the estimator suits countries with no inventory at all. Split-conformal Monte-Carlo-dropout intervals reach 81.9% coverage at a 90% nominal level, the shortfall traced to non-exchangeability under grouped splits. In conventional temporal forecasting, learners beat naive persistence only marginally and inconsistently across evaluation windows. Data and code are open.

DOI

https://doi.org/10.31223/X5G80H

Subjects

Environmental Sciences

Keywords

CO2 emissions, hybrid deep learning, Kaya identity, clustered inference, conformal prediction, explainable artificial intelligence, greenhouse gas inventories, Africa

Dates

Published: 2026-09-09 15:55

Last Updated: 2026-09-09 15:55

License

CC BY Attribution 4.0 International

Additional Metadata

Conflict of interest statement:
Author SCA is affiliated with Premium Intelligence Networks and Systems Co. Ltd. (PINSGH), a commercial company. PINSGH provided no funding or other support for this study, and there are no patents, products in development, or marketed products associated with this research to declare. This does not alter our adherence to PLOS policies on sharing data and materials. The remaining authors have declared that no competing interests exist.

Metrics

Views: 11

Downloads: 0