Pith. sign in

REVIEW 4 major objections 6 minor 83 references

RAINER: A Robust Ensemble Learning Grid Search-Tuned Framework for Rainfall Patterns Prediction

T0 review · 4 major / 6 minor · reviewed 2026-08-10 · deepseek-v4-flash

Pith's one-line read Grid-searched random forest hits 93.8 percent accuracy on next-day rain classification.

desk verdict Broad but leaky benchmark; headline results are not out-of-sample. read the letter →

arxiv 2501.16900 v1 pith:XT2HLRFO submitted 2025-01-28 cs.LG

classification cs.LG
keywords rainfallpredictionfeatureengineeringgridsearchensemblelearningrandomforestprincipalcomponentanalysisvotingclassifierKolmogorov-ArnoldNetworks
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

The paper proposes RAINER, a rainfall-prediction pipeline that combines outlier capping, missing-value imputation, engineered features such as temperature and humidity differences, PCA-based dimensionality reduction, grid-search hyperparameter tuning, and voting ensembles across weak and deep classifiers. Its central claim is that this systematic pipeline achieves top results on the Australian Bureau of Meteorology dataset, with a tuned Random Forest reaching 93.8 percent accuracy and 93.8 percent AUC for next-day rain classification. The authors argue that for structured, binary weather data, carefully preprocessed simple classifiers outperform newer deep architectures such as transformers and Kolmogorov-Arnold Networks.

What carries the argument

The load-bearing mechanism is the preprocessing pipeline applied before any model is trained: high-missingness columns are dropped, outliers in Rainfall, WindSpeed9am and WindSpeed3pm are capped, missing values are mean- or mode-imputed, and new features are constructed as the maximum difference between daily temperature and humidity readings, with the original columns removed. PCA then reduces the 17-dimensional feature space and the dataset is balanced by randomly under-sampling negative labels. Grid search, guided by a preliminary ROC-based exploration of parameter ranges, tunes every classifier, and voting ensembles combine three models at once.

What would settle it

Build the identical pipeline but confine outlier thresholds, feature-construction choices, PCA transformation, and under-sampling to the training fold alone, then evaluate on the untouched test fold; if accuracy and AUC fall materially below 93.8 percent, the paper's core claim depends on test information leaking into the model-building process.

Watch

Extended reading notes

Core claim

On the paper's own terms, the discovery is that the combination of feature construction and systematic grid search lets a Random Forest classifier beat every other method on the benchmark, including neural-network ensembles. The reported numbers put RF at 93.8 percent accuracy and 93.8 percent AUC under the 'Selected + Constructed Features' strategy, with voting ensembles like DT+LR+RF achieving a 97.3 percent recall alongside high precision.

Load-bearing premise

Every preprocessing and feature-selection decision, including which features are most correlated with next-day rain, the PCA fit, and the class balancing, is computed on the full dataset before the data is split into training and test sets.

Editorial extensions

If this is right

  • If the reported numbers hold, a tuned Random Forest is the best published next-day rain classifier on this dataset.
  • The results imply that domain-informed constructed features preserve predictive information that PCA-only reductions discard.
  • They also imply that deep sequence models and Kolmogorov-Arnold Networks add computation without accuracy gains on this structured tabular task.
  • The pipeline's voting ensembles show that precision-recall trade-offs can be stabilised by combining divergent classifiers.
  • The feature-engineering recipe could transfer to other binary tabular forecasting tasks.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • The headline figures likely depend on target-informed preprocessing being done before the 8:1:1 split, so a fully nested pipeline that learns thresholds, PCA, and balancing only from the training fold would probably report lower test metrics.
  • A natural test is to apply the same pipeline to other weather datasets with stronger seasonal or geographic structure, where the constructed 'difference' features may matter more or less.
  • The comparison would be stronger with an explicit baseline of the same models trained on raw features alone, without any target-aware construction.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

4 major / 6 minor

Summary. The paper presents RAINER, a rainfall-prediction framework for the Australian Bureau of Meteorology dataset. The pipeline removes high-missingness features, imputes missing values, caps outliers, constructs temperature/humidity difference features, balances the class distribution, applies PCA, and then trains a broad set of classifiers and neural networks with grid-search tuning and voting ensembles. The main empirical claim, stated in the abstract and Section 6.2, is that this framework achieves state-of-the-art results, with Random Forest reaching 93.8% accuracy and 93.8% AUC under the 'Selected + Constructed Features' strategy.

Significance. The paper has useful breadth: it compares many models, including recent architectures such as KAN, and systematically explores feature-engineering strategies and hyperparameter settings on a well-known public dataset. If the evaluation were valid, the model comparison and the emphasis on grid-search tuning and ensemble voting would be a useful reference for practitioners. However, the load-bearing evaluation protocol is not a valid out-of-sample assessment: class balancing, feature selection, and PCA are performed on the full dataset before the train/test split, so the test labels and test feature values influence the reported metrics. In addition, no external or published baseline is used, so the state-of-the-art claim is unsupported. The reported numbers in Table 3 cannot currently be interpreted as honest estimates of predictive performance.

major comments (4)
  1. [Section 4.4.1 and Section 5] The class-balancing step is applied to the full dataset before the 8:1:1 split. The text states that 'all positive samples are preserved, and an equal number of negative samples are randomly selected from the remaining data.' Because this selection uses RainTomorrow labels from the entire dataset, the eventual test fold is not independent: its class ratio is forced to 50/50, and its composition depends on labels that should be unseen. Every test metric in Table 3, including the headline 93.8% accuracy and AUC of Random Forest under 'Selected + Constructed Features', is therefore not a valid out-of-sample estimate. The protocol must split first, then balance only the training portion, and evaluate on an untouched test set.
  2. [Sections 4.2, 4.4, and 4.4.2] Feature selection and dimensionality reduction use the full dataset. The correlation weights in Figure 13 are computed against RainTomorrow on all records and are used to justify retaining and constructing features; PCA is fit on the full and on the full balanced dataset, and the PCA-based feature strategies in Table 3 are evaluated on components learned from data that include the test fold. The mean/mode imputation statistics and outlier-capping thresholds in Sections 4.1.1 and 4.1.2 also appear to be estimated from the full dataset. All such data-dependent steps must be learned from the training fold only and then applied to the test fold; otherwise test information leaks into feature and model selection.
  3. [Abstract and Section 6.2] The claim of 'state-of-the-art results' is unsupported by the experiments. Table 3 compares only models run in this paper; there is no comparison with previously published results on the same Australian Bureau of Meteorology dataset, no external baseline, and no defined protocol (such as repeated runs with variance or a fixed public benchmark) against which 'state-of-the-art' can be judged. A valid comparison to published work is needed before any SOTA claim can be made.
  4. [Section 5] The 8:1:1 split ratio is selected based on the authors' 'pre-exploration' experiments comparing different split ratios. If the same test fold was used to choose this ratio, then the test set has been used for model selection. The paper should clarify whether the comparison was performed on a separate validation set and, if not, use nested validation or a fixed hold-out that is never used for any design choice.
minor comments (6)
  1. [Sections 3 and 4.4] The record count is inconsistent: Section 3 states 145,460 records, while Section 4.4 states n = 200,000 daily records. The number of features is also given as 19 in Section 4.1 and as 17 in Section 4.4; these figures should be reconciled.
  2. [Section 4.2] Equations (1), (2), and (3) are not connected to the analysis: Equation (1) is a simple linear regression, Equation (2) is Bayes' rule, and Equation (3) is a set notation for a dataset. None is used or tested; either use them substantively or remove them.
  3. [Figure 13] The term 'correlation weights' is not defined. The paper should state how the values (e.g., 13.44% for MaxDifferenceTemp) are computed and how they support the feature-selection decision.
  4. [Section 4.1.2] The outlier-capping thresholds (Rainfall above 3.2 mm, WindSpeed9am above 55 km/h, WindSpeed3pm above 57 km/h) are given without justification. For Rainfall, a cap of 3.2 mm appears to remove much of the dynamic range relevant to the target variable; the choice should be justified or tested.
  5. [Table 3] Table 3 contains stray markup characters such as colons and dotted underlines and is very hard to read; it should be regenerated with clean formatting.
  6. [Figures 10 and 11] The axis labels in Figures 10 and 11 contain garbled Unicode tokens such as '/uni00000014'; the figures should be regenerated.

Circularity Check

2 steps flagged · score 5.0 of 10

Test-set leakage from full-data balancing, PCA fitting, and target-correlation feature selection makes the reported state-of-the-art metrics partially circular.

  1. fitted input called prediction [Section 4.2, Figure 13 paragraph and Table 3]
    "To further evaluate the correlation between various features and the likelihood of rain the next day, we created a pie chart in Figure 13 that depicts the correlation weights of different features. The analysis shows that MaxDifferenceTemp and Rainfall are the most correlated features with RainTomorrow, accounting for 13.44% and 13.06% of the correlation weight, respectively. This supports the decisions made in our feature selection and reconstruction process."

    The 'Selected + Constructed Features' strategy is justified by correlations with RainTomorrow computed on the full dataset before the 8:1:1 split introduced in Section 5. RainTomorrow is exactly the target later predicted on the test fold in Table 3. Thus the feature set used to report the headline 93.8% accuracy and AUC has been selected using the test labels; the test evaluation is not independent of the feature-selection input. The reported 'prediction' is partly a report on a feature space fitted to the target.

  2. fitted input called prediction [Section 4.4.1, Section 4.4.2, Section 5, Table 3]
    "However, the original dataset is imbalanced, as it contains more samples with negative labels than positive ones. To balance the dataset, all positive samples are preserved, and an equal number of negative samples are randomly selected from the remaining data. This process ensures an even distribution of positive and negative labels. ... we observed that the 8:1:1 dataset split ratio consistently outperformed others (e.g., 6:2:2 and 4:3:3). Therefore, we adopted the 8:1:1 split ratio as the standard configuration for all subsequent experiments."

    Balancing is performed on the full dataset in Section 4.4.1, and PCA is fit on the balanced dataset in Section 4.4.2, both before the train/validation/test split is defined in Section 5. The test fold is therefore a subset of the data used to subsample negatives and estimate PCA components; the test labels and test feature values are present during preprocessing. Table 3 then reports 'test' metrics for models trained on this already test-informed representation. The evaluation is circular in the sense that the construction of the test fold itself depends on information from that same fold.

full rationale

This paper is an empirical benchmark rather than a formal derivation, so most derivation-circularity patterns do not apply: there is no load-bearing self-citation chain, no uniqueness theorem imported from the authors, and no ansatz smuggled in via citation. However, the central claim of state-of-the-art performance rests on test metrics whose feature space, class balance, and PCA representation were constructed using the full dataset, including the records that later form the test fold. The feature-selection pie chart (Figure 13) explicitly uses RainTomorrow correlations on all data to justify the constructed features, and Section 4.4.1 balances the full dataset before the Section 5 split. Consequently, the reported 93.8% accuracy and 93.8% AUC for Random Forest under 'Selected + Constructed Features' are not clean out-of-sample estimates; the evaluation is partially circular because the test data influenced the preprocessing that produced the final model inputs. This warrants a score of 5, reflecting substantial leakage without a full by-construction equivalence.

Assumptions & free parameters 4 free parameters · 3 assumptions · 0 invented entities

The framework introduces no new physical entities, but it relies on several hand-chosen thresholds (outlier caps, missingness cutoff, balancing ratio, PCA component counts) and on strong assumptions about the validity of random splitting of time-series data and the safety of imputation. The absence of any sensitivity analysis for these choices is a significant weakness.

free parameters (4)
  • Outlier capping thresholds = Rainfall=3.2 mm; WindSpeed9am=55 km/h; WindSpeed3pm=57 km/h
    Chosen by visual inspection of boxplots (Section 4.1.2); these thresholds change the distribution of features used by all models.
  • Missingness removal cutoff = about 30% (features with >30% missing removed)
    Evaporation, Sunshine, Cloud9am, Cloud3pm removed; cutoff not explicitly stated, inferred from text.
  • Class balancing ratio = 1:1 negative:positive via random under-sampling
    All positive samples kept, equal number of negatives randomly selected (Section 4.4.1); changes the prior distribution and inflates accuracy.
  • PCA component counts = 2, 8, and 13 components used in feature strategies
    Selected + Constructed + PC1+PC2 and +PC1-PC8; number of components chosen by variance threshold (13 for 98.8%) and by strategy.
assumptions (3)
  • domain assumption Random 8:1:1 split of daily weather records is valid despite temporal autocorrelation
    The dataset is a 10-year time series across stations, but the paper splits randomly and does not control for temporal leakage; surrounding days can appear in both train and test.
  • domain assumption Mean and mode imputation do not distort the relationship between features and RainTomorrow
    Imputation is applied before modeling (Section 4.1.1) and treated as harmless.
  • domain assumption The BoM RainTomorrow labels are ground truth
    No external validation or error analysis of the labels.

how reviews work

0 comments
Cite this review

Pith. "Pith review of RAINER: A Robust Ensemble Learning Grid Search-Tuned Framework for Rainfall Patterns Prediction." pith.science (2026). https://pith.science/paper/XT2HLRFO

@misc{pith2026250116900,
  author       = {Pith},
  title        = {Pith review of: RAINER: A Robust Ensemble Learning Grid Search-Tuned Framework for Rainfall Patterns Prediction},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/XT2HLRFO}},
  note         = {Machine review of arXiv:2501.16900}
}
read the original abstract

Rainfall prediction remains a persistent challenge due to the highly nonlinear and complex nature of meteorological data. Existing approaches lack systematic utilization of grid search for optimal hyperparameter tuning, relying instead on heuristic or manual selection, frequently resulting in sub-optimal results. Additionally, these methods rarely incorporate newly constructed meteorological features such as differences between temperature and humidity to capture critical weather dynamics. Furthermore, there is a lack of systematic evaluation of ensemble learning techniques and limited exploration of diverse advanced models introduced in the past one or two years. To address these limitations, we propose a robust ensemble learning grid search-tuned framework (RAINER) for rainfall prediction. RAINER incorporates a comprehensive feature engineering pipeline, including outlier removal, imputation of missing values, feature reconstruction, and dimensionality reduction via Principal Component Analysis (PCA). The framework integrates novel meteorological features to capture dynamic weather patterns and systematically evaluates non-learning mathematical-based methods and a variety of machine learning models, from weak classifiers to advanced neural networks such as Kolmogorov-Arnold Networks (KAN). By leveraging grid search for hyperparameter tuning and ensemble voting techniques, RAINER achieves promising results within real-world datasets.

Figures

Figures reproduced from arXiv: 2501.16900 by the authors.

Figure 1
Figure 1. Bar Chart of missing values for meteorological features. Features such as "Evaporation" and "Cloud9am" show the highest missingness, while others like [PITH_FULL_IMAGE:figures/full_fig_p005_1.png] view at source ↗
Figure 2
Figure 2. Weather attribute dendrogram. It highlights relationships such as the strong clustering of "Pressure9am" and "Pressure3pm. [PITH_FULL_IMAGE:figures/full_fig_p006_2.png] view at source ↗
Figure 3
Figure 3. Weather feature correlation matrix. The correlation heatmap illustrates the dependencies between features, revealing strong correlations such as those between [PITH_FULL_IMAGE:figures/full_fig_p007_3.png] view at source ↗
Figures from the paper (15 more)
Figure 4
Figure 4. Figure 4: Histograms comparing feature distributions before and after imputation. The x-axis represents various meteorological attributes and the y-axis indicates the [PITH_FULL_IMAGE:figures/full_fig_p008_4.png]
Figure 5
Figure 5. Figure 5: Missing value matrix. The missing value matrix displays the distribution [PITH_FULL_IMAGE:figures/full_fig_p008_5.png]
Figure 7
Figure 7. Figure 7: Boxplots of numerical features before outlier handling. [PITH_FULL_IMAGE:figures/full_fig_p009_7.png]
Figure 8
Figure 8. Figure 8: Boxplots of numerical features after outlier handling. [PITH_FULL_IMAGE:figures/full_fig_p009_8.png]
Figure 9
Figure 9. Figure 9: Pair plot of features. The pair plot visualizes pairwise relationships among the first 950 records of selected features, providing insights into feature correlations, [PITH_FULL_IMAGE:figures/full_fig_p010_9.png]
Figure 10
Figure 10. Figure 10: Bar chart of average rainfall by month. The bar chart shows the average monthly rainfall, highlighting significant seasonal variations. June, July, and August [PITH_FULL_IMAGE:figures/full_fig_p011_10.png]
Figure 11
Figure 11. Figure 11: Bar chart of average rainfall by location. The bar chart illustrates the average rainfall for different locations. Locations 18, 33, and 43 exhibit significantly [PITH_FULL_IMAGE:figures/full_fig_p011_11.png]
Figure 12
Figure 12. Figure 12: Line plots of relationships between meteorological features. The line plots depict the relationships between various meteorological features and the likelihood [PITH_FULL_IMAGE:figures/full_fig_p012_12.png]
Figure 13
Figure 13. Figure 13: Pie chart of feature correlations with "RainTomorrow". The pie chart illustrates the correlation weights of different meteorological features with "RainTo [PITH_FULL_IMAGE:figures/full_fig_p013_13.png]
Figure 14
Figure 14. Figure 14: Correlation heatmap of features. The heatmap visualizes the strength of correlations between features, with darker colors representing stronger correlations [PITH_FULL_IMAGE:figures/full_fig_p014_14.png]
Figure 15
Figure 15. Figure 15: t-SNE visualization of weather data. The t-SNE plot shows distinct clusters for the target label categories "0" (no rain) and "1" (rain), confirming the [PITH_FULL_IMAGE:figures/full_fig_p015_15.png]
Figure 16
Figure 16. Figure 16: Percentage of variance explained by different principal components. We show the scree plot for original and balanced datasets. [PITH_FULL_IMAGE:figures/full_fig_p016_16.png]
Figure 17
Figure 17. Figure 17: Comparison of biplot representations for the original (left) and balanced (right) datasets. The biplots illustrate the contributions of meteorological features to [PITH_FULL_IMAGE:figures/full_fig_p017_17.png]
Figure 18
Figure 18. Figure 18: PCA subspace representation of observations for the first two principal components. For the original and balanced datasets, color-coded by cos2 values are [PITH_FULL_IMAGE:figures/full_fig_p018_18.png]
Figure 19
Figure 19. Figure 19: ROC curves for weak/advanced ML models under different specific model settings and training settings (e.g., 8:1:1, 6:2:2, and 4:3:3 dataset split ratios), [PITH_FULL_IMAGE:figures/full_fig_p019_19.png]

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

83 extracted references · 64 canonical work pages

  1. [1]

    Rahman, N

    F. Rahman, N. Finkelstein, A. Alyakin, N. A. Gilotra, J. Trost, S. P. Schulman, S. Saria, Using Machine Learning for Early Prediction of Cardiogenic Shock in Patients With Acute Heart Failure, Journal of the Society for Cardiovascular Angiography & Interventions 1 (3) (2022) 100308. doi:https://doi.org/10.1016/j.jscai.2022.100308

  2. [2]

    Alanazi, Using machine learning for healthcare challenges and opportunities, Informatics in Medicine Unlocked 30 (2022) 100924

    A. Alanazi, Using machine learning for healthcare challenges and opportunities, Informatics in Medicine Unlocked 30 (2022) 100924. doi:https://doi.org/10.1016/j.imu.2022.100924

  3. [3]

    Ricciardi, A

    C. Ricciardi, A. M. Ponsiglione, A. Scala, A. Borrelli, M. Misasi, G. Romano, G. Russo, M. Triassi, G. Improta, Machine Learning and Regression Analysis to Model the Length of Hospital Stay in Patients with Femur Fracture, Bioengineering 9 (2022). doi:10.3390/bioengineering9040172

  4. [4]

    S. Dev, H. Wang, C. S. Nwosu, N. Jain, B. Veeravalli, D. John, A predictive analytics approach for stroke prediction using machine learning and neural networks, Healthcare Analytics 2 (2022) 100032. doi:https://doi.org/10.1016/j. health.2022.100032

  5. [5]

    M. M. Talha, H. U. Khan, S. Iqbal, M. Alghobiri, T. Iqbal, M. Fayyaz, Deep learning in news recommender systems: A compre- hensive survey, challenges and future trends, Neurocomputing 562 (2023) 126881. doi:https://doi.org/10.1016/j. neucom.2023.126881

  6. [6]

    H. Wang, Y . Li, K. Gong, M. S. Pathan, S. Xi, B. Zhu, Z. Wen, S. Dev, MFCSNet: A Musician–Follower Complex Social Network for Measuring Musical Influence, Entertainment Computing 48 (2024) 100601. doi:https://doi.org/10. 1016/j.entcom.2023.100601. 23

  7. [7]

    J. Xu, Z. Chen, J. Li, S. Yang, H. Wang, E. C. Ngai, AlignGroup: Learning and Aligning Group Consensus with Member Pref- erences for Group Recommendation, in: Proceedings of the 33rd ACM International Conference on Information and Knowledge Management, 2024, pp. 2682–2691. doi:https://doi.org/10.1145/3627673.3679697

  8. [8]

    J. Xu, Z. Chen, S. Yang, J. Li, H. Wang, E. C.-H. Ngai, MENTOR: Multi-level Self-supervised Learning for Multimodal Recom- mendation, arXiv preprint arXiv:2402.19407 (2024). doi:https://doi.org/10.48550/arXiv.2402.19407

Show all 83 references
  1. [9]

    J. Xu, Z. Chen, J. Li, S. Yang, W. Wang, X. Hu, E. C.-H. Ngai, FourierKAN-GCF: Fourier Kolmogorov-Arnold Network–An Effective and Efficient Feature Transformation for Graph Collaborative Filtering, arXiv preprint arXiv:2406.01034 (2024)

  2. [10]

    Prasad, R

    R. Prasad, R. C. Deo, Y . Li, T. Maraseni, Soil moisture forecasting by a hybrid machine learning technique: ELM integrated with ensemble empirical mode decomposition, Geoderma 330 (2018) 136–161. doi:https://doi.org/10.1016/j. geoderma.2018.05.035

  3. [11]

    Rezaei, M

    M. Rezaei, M. A. Moghaddam, J. Piri, G. Azizyan, A. A. Shamsipour, Drought prediction using advanced hybrid machine learning for arid and semi-arid environments, KSCE Journal of Civil Engineering (2024). doi:https://doi.org/10. 1016/j.kscej.2024.100025

  4. [12]

    F. Meng, T. Ren, Z. Liu, Z. Zhong, Toward earthquake early warning: A convolutional neural network for rapid earthquake mag- nitude estimation, Artificial Intelligence in Geosciences 4 (2023) 39–46. doi:https://doi.org/10.1016/j.aiig. 2023.03.001

  5. [13]

    K. Ng, Y . Huang, C. Koo, K. Chong, A. El-Shafie, A. Najah Ahmed, A review of hybrid deep learning applications for streamflow forecasting, Journal of Hydrology 625 (2023) 130141.doi:https://doi.org/10.1016/j.jhydrol.2023.130141

  6. [14]

    L. Ma, Y . Liu, X. Zhang, Y . Ye, G. Yin, B. A. Johnson, Deep learning in remote sensing applications: A meta-analysis and review, ISPRS Journal of Photogrammetry and Remote Sensing 152 (2019) 166–177. doi:https://doi.org/10.1016/ j.isprsjprs.2019.04.015

  7. [15]

    Schulz, R

    K. Schulz, R. Hänsch, U. Sörgel, Machine learning methods for remote sensing applications: an overview, in: ISPRS international workshop on machine learning and remote sensing/GIS applications IX, 2018. doi:https://doi.org/10.1117/12. 2503653

  8. [16]

    Talukdar, P

    S. Talukdar, P. Singha, S. Mahato, S. Pal, Y . Liou, Land-use land-cover classification by machine learning classifiers for satellite observations—A review, Remote Sensing 12 (7) (2020) 1135. doi:https://doi.org/10.3390/rs12071135

  9. [17]

    K. Cui, R. Li, S. L. Polk, Y . Lin, H. Zhang, J. M. Murphy, R. J. Plemmons, R. H. Chan, Superpixel-based and Spatially-regularized Diffusion Learning for Unsupervised Hyperspectral Image Clustering, IEEE Transactions on Geoscience and Remote Sensing (2024)

  10. [18]

    K. Cui, Z. Shao, G. Larsen, V . Pauca, S. Alqahtani, D. Segurado, J. a. Pinheiro, M. Wang, D. Lutz, R. Plemmons, M. Silman, PalmProbNet: A Probabilistic Approach to Understanding Palm Distributions in Ecuadorian Tropical Forest via Transfer Learn- ing, in: Proceedings of the 2...

  11. [19]

    Y . Li, H. Wang, J. Xu, Z. Ma, P. Wu, S. Wang, S. Dev, CP2M: Clustered-Patch-Mixed Mosaic Augmentation for Aerial Image Segmentation, arXiv preprint arXiv:2501.15389 (2025). doi:https://doi.org/10.48550/arXiv.2501.15389

  12. [20]

    S. Dev, A. Nautiyal, Y . H. Lee, S. Winkler, CloudSegNet: A Deep Network for Nychthemeron Cloud Image Segmentation, IEEE Geoscience and Remote Sensing Letters 16 (12) (2019) 1814–1818. doi:10.1109/LGRS.2019.2912140

  13. [21]

    Y . Li, H. Wang, S. Wang, Y . H. Lee, M. Salman Pathan, S. Dev, UCloudNet: A Residual U-Net with Deep Supervision for Cloud Image Segmentation, in: IGARSS 2024 - 2024 IEEE International Geoscience and Remote Sensing Symposium, 2024, pp. 5553–5557. doi:10.1109/IGARSS53475.2024.10640450

  14. [22]

    Y . Li, H. Wang, J. Xu, P. Wu, Y . Xiao, S. Wang, S. Dev, DDUNet: Dual Dynamic U-Net for Highly-Efficient Cloud Segmentation, arXiv preprint arXiv:2501.15385 (2025). doi:https://doi.org/10.48550/arXiv.2501.15385

  15. [23]

    Abebe, H

    E. Abebe, H. Kebede, M. Kevin, Z. Demissie, Earthquakes magnitude prediction using deep learning for the Horn of Africa, Soil Dynamics and Earthquake Engineering 170 (2023) 107913.doi:https://doi.org/10.1016/j.soildyn.2023. 107913

  16. [24]

    Ortega, Machine Learning and Seismic Hazard: A Combination of Probabilistic Approaches for Probabilistic Seismic Hazard Analysis (2024)

    R. Ortega, Machine Learning and Seismic Hazard: A Combination of Probabilistic Approaches for Probabilistic Seismic Hazard Analysis (2024). doi:10.5772/intechopen.1006533

  17. [25]

    F. T. Teshome, H. K. Bayabil, B. Schaffer, Y . Ampatzidis, G. Hoogenboom, Improving soil moisture prediction with deep learning and machine learning models, Computers and Electronics in Agriculture 226 (2024) 109414. doi:https://doi.org/10. 1016/j.compag.2024.109414

  18. [26]

    Y . Wang, L. Shi, Y . Hu, X. Hu, W. Song, L. Wang, A comprehensive study of deep learning for soil moisture prediction, Hydrology and Earth System Sciences 28 (4) (2024) 917–943. doi:10.5194/hess-28-917-2024

  19. [27]

    Márquez-Grajales, R

    A. Márquez-Grajales, R. Villegas-Vega, F. Salas-Martínez, H.-G. Acosta-Mesa, E. Mezura-Montes, Characterizing drought pre- diction with deep learning: A literature review, MethodsX 13 (2024) 102800.doi:https://doi.org/10.1016/j.mex. 2024.102800

  20. [28]

    Nandgude, T

    N. Nandgude, T. P. Singh, S. Nandgude, M. Tiwari, Drought prediction: A comprehensive review of different drought prediction models and adopted technologies, Sustainability 15 (15) (2023). doi:10.3390/su151511684

  21. [29]

    Anilkumar, R

    R. Anilkumar, R. Bharti, D. Chutia, S. P. Aggarwal, Modelling point mass balance for the glaciers of the central european alps using machine learning techniques, The Cryosphere 17 (7) (2023) 2811–2828. doi:10.5194/tc-17-2811-2023

  22. [30]

    Li, C.-Y

    W. Li, C.-Y . Hsu, M. Tedesco, Advancing arctic sea ice remote sensing with ai and deep learning: Opportunities and challenges, Remote Sensing 16 (20) (2024). doi:10.3390/rs16203764

  23. [31]

    H. Wang, M. S. Pathan, Y . H. Lee, S. Dev, Day-ahead Forecasts of Air Temperature, in: 2021 IEEE USNC-URSI Radio Science Meeting (Joint with AP-S Symposium), 2021, pp. 94–95. doi:10.23919/USNC-URSI51813.2021.9703606

  24. [32]

    Cifuentes, G

    J. Cifuentes, G. Marulanda, A. Bello, J. Reneses, Air temperature forecasting using machine learning techniques: A review, Energies 13 (16) (2020). doi:10.3390/en13164215. 25

  25. [33]

    M. S. Rahman, F. A. Tumpa, M. S. Islam, A. A. Arabi, M. S. B. Hossain, M. S. U. Haque, Comparative evaluation of weather forecasting using machine learning models, 2023, pp. 1–6. doi:10.1109/ICCIT60459.2023.10441077

  26. [34]

    K. Cui, W. Tang, R. Zhu, M. Wang, G. D. Larsen, V . P. Pauca, S. Alqahtani, F. Yang, D. Segurado, P. Fine, et al., Real-Time Localization and Bimodal Point Pattern Analysis of Palms Using UA V Imagery, arXiv preprint arXiv:2410.11124 (2024)

  27. [35]

    M. M. Hassan, M. A. T. Rony, M. A. R. Khan, M. M. Hassan, F. Yasmin, A. Nag, T. H. Zarin, A. K. Bairagi, S. Alshathri, W. El-Shafai, Machine Learning-Based Rainfall Prediction: Unveiling Insights and Forecasting for Improved Preparedness, IEEE Access 11 (2023) 132196–132222. d...

  28. [36]

    S. H. Pour, S. Shahid, E.-S. Chung, A Hybrid Model for Statistical Downscaling of Daily Rainfall, Procedia Engineering 154 (2016) 1424–1430, 12th International Conference on Hydroinformatics (HIC 2016) - Smart Water for the Future.doi:https: //doi.org/10.1016/j.proeng.2016.07.514

  29. [37]

    P. Das, D. A. Sachindra, K. Chanda, Machine Learning-Based Rainfall Forecasting with Multiple Non-Linear Feature Selection Algorithms, Water Resources Management 36 (2022) 6043–6071. doi:10.1007/s11269-022-03341-8

  30. [38]

    V . A. Martinez Lopez, G. van Urk, P. J. Doodkorte, M. Zeman, O. Isabella, H. Ziar, Using sky-classification to improve the short-term prediction of irradiance with sky images and convolutional neural networks, Solar Energy 269 (2024) 112320. doi: https://doi.org/10.1016/j.sol...

  31. [39]

    M. S. Pathan, A. Nag, S. Dev, Efficient rainfall prediction using a dimensionality reduction method, in: IGARSS 2022 - 2022 IEEE International Geoscience and Remote Sensing Symposium, 2022, pp. 6737–6740. doi:10.1109/IGARSS46834. 2022.9884849

  32. [40]

    Kustiyo, A

    A. Kustiyo, A. Buono, A. Faqih, K. Priandana, Analysis on dimensionality reduction techniques for sub-seasonal to seasonal rainfall prediction, in: 2021 International Conference on Computer System, Information Technology, and Electrical Engineering (COSITE), 2021, pp. 156–160....

  33. [41]

    Y . Meng, S. N. Qasem, M. Shokri, S. S, Dimension reduction of machine learning-based forecasting models employing principal component analysis, Mathematics 8 (8) (2020). doi:10.3390/math8081233

  34. [42]

    R. Qiu, C. Liu, N. Cui, Y . Gao, L. Li, Z. Wu, S. Jiang, M. Hu, Generalized extreme gradient boosting model for predicting daily global solar radiation for locations without historical data, Energy Conversion and Management 258 (2022) 115488. doi: 10.1016/j.enconman.2022.115488

  35. [43]

    E. K. Sahin, Assessing the predictive capability of ensemble tree methods for landslide susceptibility mapping using xgboost, gra- dient boosting machine, and random forest, SN Applied Sciences 2 (8) (2020) 1308.doi:10.1007/s42452-020-3060-1

  36. [44]

    G. Rong, S. Alu, K. Li, Y . Su, J. Zhang, Y . Zhang, T. Li, Rainfall Induced Landslide Susceptibility Mapping Based on Bayesian Optimized Random Forest and Gradient Boosting Decision Tree Models—A Case Study of Shuicheng County, China, Water 12 (11) (2020). doi:10.3390/w12113066. 26

  37. [45]

    Z. Liu, Y . Wang, S. Vaidya, F. Ruehle, J. Halverson, M. Solja ˇci´c, T. Y . Hou, M. Tegmark, Kan: Kolmogorov-arnold networks, arXiv preprint arXiv:2404.19756 (2024). doi:10.48550/arXiv.2404.19756

  38. [46]

    Bauer, A

    P. Bauer, A. Thorpe, G. Brunet, The quiet revolution of numerical weather prediction, Nature 525 (2015) 47–55. doi:10. 1038/nature14956

  39. [47]

    Steiner, J

    M. Steiner, J. A. Smith, S. J. Burges, Climatological characteristics of three-dimensional storm structure from operational radar and rain gauge data, Journal of applied meteorology 34 (9) (1995) 1978–2007. doi:10.1175/1520-0450(1995) 034<1978:CCOTDS>2.0.CO;2

  40. [48]

    E. Matricciani, A mathematical theory of de-integrating long-time integrated rainfall and its application for predicting 1-min rain rate statistics, International Journal of Satellite Communications and Networking 29 (6) (2011) 443–457.doi:10.1002/sat. 990

  41. [49]

    K. Lee, S. Pyo, M. Wong, Spatiotemporal aerosol prediction model based on fusion of machine learning and spatial analysis, Asian Journal of Atmospheric Environment (2024). doi:10.1007/s44273-024-00031-2

  42. [50]

    C. Gao, X. Guan, M. J. Booij, Y . Meng, Y .-P. Xu, A new framework for a multi-site stochastic daily rainfall model: Coupling a univariate Markov chain model with a multi-site rainfall event model, Journal of Hydrology 598 (2021) 126478. doi:10. 1016/j.jhydrol.2021.126478

  43. [51]

    Usman Saeed Khan, K

    M. Usman Saeed Khan, K. Mohammad Saifullah, A. Hussain, H. Mohammad Azamathulla, Comparative analysis of different rainfall prediction models: A case study of Aligarh City, India, Results in Engineering 22 (2024) 102093. doi:https: //doi.org/10.1016/j.rineng.2024.102093

  44. [52]

    V . B. Nikam, B. Meshram, Modeling Rainfall Prediction Using Data Mining Method: A Bayesian Approach, in: 2013 Fifth Inter- national Conference on Computational Intelligence, Modelling and Simulation, 2013, pp. 132–136. doi:10.1109/CIMSim. 2013.29

  45. [53]

    Permanasari, I

    A. Permanasari, I. Hidayah, I. A. Bustoni, SARIMA (Seasonal ARIMA) implementation on time series to forecast the number of Malaria incidence, 2013, pp. 203–207. doi:10.1109/ICITEED.2013.6676239

  46. [54]

    A. Sani, A. Muhammad Auwal, M. Adenomon, Application of Sarima Models in Modelling and Forecasting Monthly Rainfall in Nigeria, Asian Journal of Probability and Statistics 13 (2021) 30–43. doi:10.9734/AJPAS/2021/v13i330310

  47. [55]

    R. L. Wilby, C. W. Dawson, E. M. Barrow, Statistical downscaling of general circulation model output: A comparison of methods, Water resources research 38 (12) (2002) 6–1. doi:10.1029/98WR02577

  48. [56]

    M.-H. Yeo, A. Frei, R. K. Gelda, E. M. Owens, A stochastic weather model for generating daily precipitation series at ungauged locations in the Catskill Mountain region of New York state, International Journal of Climatology 39 (2019) 3467–3483. doi: 10.1002/joc.6230

  49. [57]

    A. Das, S. Bhan, M. Mohapatra, Feasibility of Model Output Statistics (MOS) for Improving the Quantitative Precipitation Forecasts of IMD GFS Model, Journal of Hydrology (2024). doi:10.1016/j.jhydrol.2024.100937. 27

  50. [58]

    A. Nikahd, Advanced of Mathematics-Statistics Methods to Radar Calibration for Rainfall Estimation; A Review, Interna- tional Journal on Recent and Innovation Trends in Computing and Communication 3 (2015) 96–105. doi:10.17762/ IJRITCC2321-8169.150121

  51. [59]

    X. Wang, W. Liu, Deep Time Series Forecasting Models: A Comprehensive Survey, Mathematics 12 (10) (2024). doi:10. 3390/math12101504

  52. [60]

    Sahour, V

    H. Sahour, V . Gholami, J. Torkaman, M. Vazifedan, S. Saeedi, Random forest and extreme gradient boosting algorithms for streamflow modeling using vessel features and tree-rings, Environmental Earth Sciences 80 (18) (2021) 747. doi:10.1007/ s12665-021-10054-5

  53. [61]

    H. Tao, S. Awadh, S. Salih, S. Shafik, Z. Yaseen, Integration of extreme gradient boosting feature selection approach with machine learning models: application of weather relative humidity prediction, Neural Computing and Applications 34 (01 2022). doi:10.1007/s00521-021-06362-3

  54. [62]

    C. Wu, K. Chau, C. Fan, Prediction of rainfall time series using modular artificial neural networks coupled with data-preprocessing techniques, Journal of Hydrology 389 (1) (2010) 146–167. doi:https://doi.org/10.1016/j.jhydrol.2010.05. 040

  55. [63]

    H. T. Pham, J. Awange, M. Kuhn, Evaluation of three feature dimension reduction techniques for machine learning-based crop yield prediction models, Sensors 22 (17) (2022). doi:10.3390/s22176609

  56. [64]

    Stefenon, L

    S. Stefenon, L. Seman, E. da Silva, E. Finardi, Hypertuned wavelet convolutional neural network with long short-term memory for time series forecasting in hydroelectric power plants, Energy (2024). doi:10.1016/j.energy.2024.125311

  57. [65]

    H. Wang, Y . Li, S. Xi, S. Wang, M. S. Pathan, S. Dev, AMDCNet: An attentional multi-directional convolutional network for stereo matching, Displays 74 (2022) 102243. doi:https://doi.org/10.1016/j.displa.2022.102243

  58. [66]

    H. Wang, M. S. Pathan, S. Dev, Stereo Matching Based on Visual Sensitive Information, in: 2021 6th International Conference on Image, Vision and Computing (ICIVC), 2021, pp. 312–316. doi:10.1109/ICIVC52351.2021.9527014

  59. [67]

    Batra, H

    S. Batra, H. Wang, A. Nag, P. Brodeur, M. Checkley, A. Klinkert, S. Dev, DMCNet: Diversified model combination network for understanding engagement from video screengrabs, Systems and Soft Computing 4 (2022) 200039

  60. [68]

    H. Wang, B. Zhu, Y . Li, K. Gong, Z. Wen, S. Wang, S. Dev, SYGNet: A SVD-YOLO based GhostNet for Real-time Driving Scene Parsing, in: 2022 IEEE International Conference on Image Processing (ICIP), 2022, pp. 2701–2705

  61. [69]

    W. Tang, K. Cui, R. H. Chan, Optimized Hard Exudate Detection with Supervised Contrastive Learning, in: 2024 IEEE Interna- tional Symposium on Biomedical Imaging (ISBI), IEEE, 2024, pp. 1–5

  62. [70]

    F. Pan, Y . Wu, K. Cui, S. Chen, Y . Li, Y . Liu, A. Shakoor, H. Zhao, B. Lu, S. Zhi, et al., Accurate detection and instance segmentation of unstained living adherent cells in differential interference contrast images, Computers in Biology and Medicine 182 (2024) 109151. 28

  63. [72]

    Y . Li, H. Wang, A. Katsaggelos, CPDR: Towards Highly-Efficient Salient Object Detection via Crossed Post-decoder Refinement, in: 35th British Machine Vision Conference 2024, BMVC 2024, Glasgow, UK, November 25-28, 2024, BMV A, 2024

  64. [73]

    Z. Li, H. Wang, Y . Li, S. Dev, G. Zuo, VGRISys: A Vision-Guided Robotic Intelligent System for Autonomous Instrument Calibration*, in: 2023 IEEE International Conference on Robotics and Biomimetics (ROBIO), 2023, pp. 1–6. doi:10.1109/ ROBIO58561.2023.10354843

  65. [74]

    M. Huo, Z. Zhang, X. Ren, X. Yang, C. Ye, AbHE: All Attention-Based Homography Estimation, IEEE Transactions on Instru- mentation and Measurement 73 (2024) 1–11

  66. [75]

    Z. Wang, B. Li, C. Wang, S. Scherer, AirShot: Efficient few-shot detection for autonomous exploration, in: IEEE/RSJ Interna- tional Conference on Intelligent Robots and Systems (IROS), 2024

  67. [76]

    X. Zhu, R. Tian, C. Xu, M. Huo, W. Zhan, M. Tomizuka, M. Ding, Fanuc manipulation: A dataset for learning-based manipulation with fanuc mate 200id robot (2023)

  68. [77]

    Wang, ONLS: OPTIMAL NOISE LEVEL SEARCH IN DIFFUSION AUTOENCODERS WITHOUT FINE-TUNING, in: The Second Tiny Papers Track at ICLR 2024, 2024

    Z. Wang, ONLS: OPTIMAL NOISE LEVEL SEARCH IN DIFFUSION AUTOENCODERS WITHOUT FINE-TUNING, in: The Second Tiny Papers Track at ICLR 2024, 2024. URL https://openreview.net/forum?id=Q8diCUHTZd

  69. [78]

    H. Lin, Y . Wang, M. Huo, C. Peng, Z. Liu, M. Tomizuka, Joint Pedestrian Trajectory Prediction through Posterior Sampling, 2024 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS) (2024) 5672–5679

  70. [79]

    M. Huo, M. Ding, C. Xu, T. Tian, X. Zhu, Y . Mu, L. Sun, M. Tomizuka, W. Zhan, Human-oriented Representation Learning for Robotic Manipulation, ArXiv abs/2310.03023 (2023)

  71. [80]

    Haji-Aghajany, W

    S. Haji-Aghajany, W. Rohm, M. Kryza, K. Smolak, Machine Learning-Based Wet Refractivity Prediction Through GNSS Tropo- sphere Tomography for Ensemble Troposphere Conditions Forecasting, IEEE Transactions on Geoscience and Remote Sensing 62 (2024) 1–18. doi:10.1109/TGRS.2024.3417487

  72. [81]

    Henderson, R

    K. Henderson, R. K. Chakrabortty, A machine learning predictive model for bushfire ignition and severity: The study of australian black summer bushfires, Decision Analytics Journal 14 (2025) 100529. doi:https://doi.org/10.1016/j.dajour. 2024.100529

  73. [82]

    Lemaitre, F

    G. Lemaitre, F. Nogueira, C. K. Aridas, Imbalanced-learn: A Python Toolbox to Tackle the Curse of Imbalanced Datasets in Machine Learning, Journal of Machine Learning Research 18 (17) (2017) 1–5

  74. [83]

    van der Maaten, G

    L. van der Maaten, G. Hinton, Visualizing Data using t-SNE, Journal of Machine Learning Research 9 (86) (2008) 2579–2605

  75. [84]

    L. W. H. Abdi, Principal component analysis, Wiley Interdiscip. Rev. Comput. Stat. (4) . 2 (2010) 433–459. doi:10.1002/ wics.101. 29

Pith tools

Reviewed August 10, 2026 · model on record in the stance chip above.