Abstract
Probabilistic forecasting of photovoltaic (PV) power generation is essential for reliable grid operation under high renewable-energy penetration. Existing approaches either separate point prediction from uncertainty estimation or use historical analogues only as external pre- or post-hoc augmentation, applying static similarity weights that fail to adapt to meteorological variability. This paper proposes Contextual retrieval-augmented tree ensemble Boosting (CrateBoost), a probabilistic forecasting framework whose central mechanism is the injection of regime-matched historical analogues directly into the natural-gradient updates of a joint mean–variance booster at every boosting iteration, so that retrieval shapes the learning trajectory itself rather than correcting it externally. The framework is organized as a multi-stage pipeline comprising: (i) data pre-processing with outlier treatment and physics-informed feature engineering; (ii) a precomputed target-supervised PowerUMAP embedding, constructed from irradiance, thermal, and electrical features and indexed with Facebook AI Similarity Search (FAISS), which provides the similarity manifold over which retrieval operates and ensures that retrieved neighbours correspond to comparable PV-conversion regimes rather than superficially similar raw weather; (iii) the core boosting loop, which jointly optimizes mean and variance heads under a Gaussian negative log-likelihood objective with causality-preserving neighbour selection, blends retrieval-augmented gradients at every iteration, and modulates analogue influence according to local clear-sky-ratio dispersion so that the model leans on parametric estimates under stable conditions and on retrieved analogues under broken-cloud regimes; and (iv) a post-hoc Ensemble Model Output Statistics (EMOS) recalibration applied on a held-out validation set to deliver calibrated marginal predictive distributions. The framework is evaluated on six years (2017–2022) of operational data from the Yulara Solar System (Australia) under a ten-fold walk-forward cross-validation protocol at 1-, 3-, and 6-h horizons. Benchmarked against eight state-of-the-art methods spanning traditional machine learning (LightGBM, XGBoost, CatBoost, and NGBoost) and deep learning architectures (TFT, LSTM, Autoformer, and N-BEATS), the proposed method achieves consistent gains across horizons, outperforming the strongest baseline at each horizon by 4 to 14% in normalized mean absolute error and by 1.5 to 15% in the continuous ranked probability score.
| Original language | English |
|---|---|
| Article number | 100855 |
| Journal | Energy and AI |
| Volume | 25 |
| DOIs | |
| State | Published - Sep 2026 |
Bibliographical note
Publisher Copyright:© 2026 The Authors. Published by Elsevier Ltd. This is an open access article under the CC BY license. http://creativecommons.org/licenses/by/4.0/
UN SDGs
This output contributes to the following UN Sustainable Development Goals (SDGs)
-
SDG 7 Affordable and Clean Energy
Keywords
- Ensemble model output statistics
- Natural gradient boosting
- Photovoltaic power forecasting
- Probabilistic forecasting
- Retrieval-augmented gradient boosting
- UMAP
ASJC Scopus subject areas
- Engineering (miscellaneous)
- General Energy
- Artificial Intelligence
Fingerprint
Dive into the research topics of 'Contextual retrieval-augmented tree ensemble boosting for multi-horizon solar PV power forecasting'. Together they form a unique fingerprint.Cite this
- APA
- Author
- BIBTEX
- Harvard
- Standard
- RIS
- Vancouver