•  
  •  
 

Abstract

Receiving rainfall predictions in dry regions such as Saudi Arabia is challenging due to limited observational stations and the temporal and spatial variability of rainfall events. This study explores the performance of nine remote sensing rainfall datasets over the Qassim region of central Saudi Arabia. High-resolution rainfall data from ground-based sources were used to evaluate the accuracy of CHIRPS, TRMM, GPM, PERSIANN, CMORPH, CFSR, TerraClimate, TerraClimate_Monthly, and ERA5 datasets. Performance metrics such as R2, RMSE, MAE, and bias were calculated to quantify the accuracy of each dataset. Analysis revealed that ERA5 generally demonstrated better accuracy than the other datasets with a mean R2 and RMSE, MAE, and bias values of 0.59, 3.91, 2.89, and 0.43, respectively. The algorithms were trained, validated, and applied to forecast daily, 5-day, 10-day, and monthly total rainfall. Root mean square error (RMSE) values decreased by up to 52%, 50%, 56%, and 45% for daily, 5-day, 10-day, and monthly rainfall forecasts, respectively. To improve the accuracy of future predictions, multivariable models were developed that incorporated mean air temperature with precipitation predictors. Five machine-learning algorithms were tested: Gradient Boosting Regressor (GBR), Histogram Gradient Boosting Regressor (HGBR), Random Forest Regressor (RFR), Extreme Gradient Boosting Regressor (XGBR), and an advanced version of XGBR. For the 70/30 train-validation split over time, the models based on boosting methods (HGBR and XGBR) demonstrated the highest predictive ability and yielded the best statistical performance (R2 = 0.90) with the lowest errors. Combining a suite of reanalysis datasets with advanced machine learning strategies enables effective rainfall prediction in arid regions that lack adequate observational data. Such a strategy facilitates rainfall estimation in ungauged catchments, enabling water resource management and flash-flood risk assessment in the desert domain.

Share

COinS