Leveraging Grid Search in Tuning Ensemble Models to Predict Boiling Points of Alkanes

Authors

DOI:

https://doi.org/10.32350/sir.102.04

Keywords:

alkanes, Boiling points, Machine Learning (ML), Random Forest (RF), XGBoost

Abstract

Predicting the Normal Boiling Points (NBPs) of alkanes is a fundamental problem in physical organic chemistry and chemical engineering. Two-stage GridSearchCV workflow was used to tune Random Forest (RF) and eXtreme Gradient Boosting (XGBoost) regressors in order to predict the NBPs of alkanes containing up to ten carbon atoms. Each alkane was encoded as a unique ten-digit numerical sequence. In Stage I, broad hyperparameter ranges were explored to identify regions of optimal model performance. In Stage II, a refined grid was constructed around the Stage I optima to pinpoint the best hyperparameter combinations with greater precision. Both models were evaluated on held-out test data using Mean Squared Error (MSE) and R² metrics. After two-stage tuning, XGBoost showed substantial improvement with R²=0.987 and an MSE=12.002. While, the RF regressor showed a modest improvement with R²=0.975 and an MSE=22.538 over Stage I baselines. SHapley Additive exPlanations (SHAP) analysis was subsequently applied to quantify the contribution of individual structural features to each model's predictions. Positional features corresponding to later carbon atoms were identified. The results demonstrated that systematic two-stage hyperparameter optimization via GridSearchCV significantly enhanced predictive accuracy and interpretability in chemical property prediction tasks

Downloads

Download data is not yet available.
0

References

1. Wiener, H. (1947). Structural determination of paraffin boiling points. Journal of the American Chemical Society, 69(1), 17–20. https://doi.org/10.1021/ja01193a005

2. Wessel, M. D., & Jurs, P. C. (1995). Prediction of normal boiling points of hydrocarbons from molecular structure. Journal of Chemical Information and Computer Sciences, 35(1), 68–76. https://doi.org/10.1021/ci00023a010

3. Katritzky, A. R., Lobanov, V. S., & Karelson, M. (1995). QSPR: The correlation and quantitative prediction of chemical and physical properties from structure. Chemical Society Reviews, 24(4), 279–287. https://doi.org/10.1039/CS9952400279

4. Hosoya, H. (1971). Topological index. A newly proposed quantity characterizing the topological nature of structural isomers of saturated hydrocarbons. Bulletin of the Chemical Society of Japan, 44(9), 2332–2339. https://doi.org/10.1246/bcsj.44.2332

5. Gutman, I. (1994). Selected properties of the Schultz molecular topological index. Journal of Chemical Information and Computer Sciences, 34(5), 1087–1089. https://doi.org/10.1021/ci00021a009

6. Schultz, H. P. (1989). Topological organic chemistry. 1. Graph theory and topological indices of alkanes. Journal of Chemical Information and Computer Sciences, 29(3), 227–228. https://doi.org/10.1021/ci00063a012

7. Cherqaoui, D., Villemin, D., Mesbah, A., Cense, J.-M., & Kvasnicka, V. (1994). Use of a neural network to determine the normal boiling points of acyclic ethers, peroxides, acetals and their sulfur analogues. Journal of the Chemical Society, Faraday Transactions, 90(14), 2015–2019. https://doi.org/10.1039/FT9949002015

8. Goll, E. S., & Jurs, P. C. (1999). Prediction of the normal boiling points of organic compounds from molecular structures with acomputational neural network model. Journal of Chemical Information and Computer Sciences, 39(6), 974–983.https://doi.org/10.1021/ci990071l.

Goll and Jurs present one of the earliest systematic benchmarks of neural network models for predicting normal boiling points of structurally diverse organic compounds. Using descriptors derived from molecular topology and geometry, and combining multilinear regression with genetic-algorithm-selected descriptors and neural networks, they achieve a root-mean-square error of 11.72 K on a test set of 91 compounds. This work establishes a quantitative performance baseline against which the ensemble models of the present study can be compared, and it highlights the trajectory from single-architecture neural networks toward the more flexible, regularized tree-based ensembles employed here.

9. Espinosa, G., Yaffe, D., Cohen, Y., Arenas, A., & Giralt, F. (2000). Neural network based quantitative structural property relations (QSPRs) for predicting boiling points of aliphatic hydrocarbons. Journal of Chemical Information and Computer Sciences, 40(3), 859–879. https://doi.org/10.1021/ci000442u

10. Espinosa, G., Arenas, A., & Giralt, F. (2001). Prediction of boiling points of organic compounds from molecular descriptors by using backpropagation neural network. In R. Carbó-Dorca, X. Gironés, & P. G. Mezey (Eds.), Fundamentals of Molecular Similarity (pp.1–10). Kluwer Academic/Plenum Publishers. https://doi.org/10.1007/978-1-4757-3273-3_1

11. Liu, B., & Karimi Nouroddin, M. (2023). Application of artificial intelligent approach to predict the normal boiling point of refrigerants. International Journal of Chemical Engineering, 2023, Article 6809569. https://doi.org/10.1155/2023/6809569

12. Qu, C., Kearsley, A. J., Schneider, B. I., Keyrouz, W., & Allison, T. C. (2022). Graph convolutional neural network applied to the prediction of normal boiling point. Journal of Molecular Graphics and Modelling, 112, Article 108149. https://doi.org/10.1016/j.jmgm.2022.108149

13. Deng, S., Su, W., & Zhao, L. (2016). A neural network for predicting normal boiling point of pure refrigerants using molecular groups and a topological index. International Journal of Refrigeration, 63, 63–71. https://doi.org/10.1016/j.ijrefrig.2015.10.025

14. Nizami, A. R., Ali, S. F., & Afzal, M. Z. (2025). Novel descriptors for the prediction of molecular properties. Open Chemistry, 23(1). https://doi.org/10.1515/chem-2025-0194

This paper by the present article’s corresponding author introduces a family of novel graph-theoretic molecular descriptors derived from weighted distance and degree matrices of molecular graphs. The descriptors are evaluated for their ability to predict boiling points and other physicochemical properties of alkanes, and they are shown to achieve competitive or superior correlation compared with established topological indices. The work provides the theoretical and descriptive foundation for the structural features used in the present study, making it the primary methodological predecessor to the machine learning modelling reported here.

15. Breiman, L. (2001). Random forests. Machine Learning, 45(1), 5–32. https://doi.org/10.1023/A:1010933404324

16. Friedman, J. H. (2001). Greedy function approximation: A gradient boosting machine. The Annals of Statistics, 29(5), 1189–1232.https://doi.org/10.1214/aos/1013203451

17. Chen, T., & Guestrin, C. (2016). XGBoost: A scalable tree boosting system. In Proceedings of the 22nd ACM SIGKDD International Conference on Knowledge Discovery and Data Mining (pp. 785–794). ACM. https://doi.org/10.1145/2939672.2939785

18. Svetnik, V., Liaw, A., Tong, C., Culberson, J. C., Sheridan, R. P., & Feuston, B. P. (2003). Random forest: A classification and regression tool for compound classification and QSAR modeling. Journal of Chemical Information and Modeling, 43(6), 1947–1958. https://doi.org/10.1021/ci034160g

19. Pedregosa, F., Varoquaux, G., Gramfort, A., Michel, V., Thirion, B., Grisel, O., Blondel, M., Prettenhofer, P., Weiss, R., Dubourg, V., Vanderplas, J., Passos, A., Cournapeau, D., Brucher, M., Perrot, M., & Duchesnay, E. (2011). Scikit-learn: Machine learning in Python. Journal of Machine Learning Research, 12, 2825–2830.

20. Lundberg, S. M., Erion, G. G., & Lee, S.-I. (2018). Consistent individualized feature attribution for tree ensembles. arXiv preprint arXiv:1802.03888.

21. Lundberg, S. M., & Lee, S.-I. (2017). A unified approach to interpreting model predictions. Advances in Neural Information Processing Systems, 30, 4765–4774.

22. Chaudhury, S., Shelke, N., Rashid, Z. M., & Sau, K. (2022). Effect of grid search and hyper parameter tuned pipeline with various classifiers and PCA for breast cancer detection. Current Signal Transduction Therapy, 17(3), 45–56. https://doi.org/10.2174/1574362417666220715105527.

This paper evaluates the combined effect of PCA-based dimensionality reduction and GridSearchCV hyperparameter tuning across multiple classifiers for breast cancer detection. The authors demonstrate consistent and statistically significant performance gains attributable specifically to the grid-search step, isolating its contribution from feature engineering effects. This finding corroborates the methodological choice in the present study of using GridSearchCV as the primary optimisation tool, and it extends the evidence base for grid-search-driven tuning to high-dimensional biomedical datasets with class-imbalance challenges.

23. Alshammari, T. (2024). Using artificial neural networks with GridSearchCV for predicting indoor temperature in a smart home. Engineering, Technology & Applied Science Research, 14(2), 13437–13443. https://doi.org/10.48084/etasr.7008.

This study demonstrates the effectiveness of GridSearchCV in optimising artificial neural network hyperparameters for regression tasks in smart building systems. By systematically searching across network architecture and training parameters, the author achieved significantly lower prediction errors for indoor temperature than baseline models. The work is directly relevant to the present article as a practical application of the same two-stage grid-search strategy applied here to ensemble regressors, and it underscores the domain-agnostic utility of GridSearchCV beyond classification contexts.

24. Chen, C.-H., Tanaka, K., & Funatsu, K. (2020). Comparison and improvement of the predictability and interpretability with ensemble learning models in QSPR applications. Journal of Cheminformatics, 12, Article 19. https://doi.org/10.1186/s13321-020-0417-9

25. Hafner J, Noh J, Reif B, Riniker S. AI-powered prediction of critical properties and boiling points: A hybrid ensemble learning and QSPR approach. J Cheminform. 2025;17:62. doi:10.1186/s13321-025-01062-9.

Hafner et al. introduce a hybrid framework that integrates QSPR descriptors with stacked ensemble learning to predict normal boiling points and critical thermodynamic properties across chemically diverse datasets. The study systematically benchmarks gradient boosting, Random Forest, and blended meta-learners, demonstrating that ensemble stacking with optimised hyperparameters outperforms single-model and classical QSPR approaches. Its findings are closely aligned with the goals of the present article, providing contemporary evidence that structured hyperparameter optimisation of tree-based ensembles represents the current state of the art in computational thermophysical property prediction.

26. Mukwembi, S., & Nyabadza, F. (2021). A new model for predicting boiling points of alkanes. Scientific Reports, 11(1), Article 24261. https://doi.org/10.1038/s41598-021-03541-z

Published

2025-06-25

How to Cite

1.
Nizami AR, Ali SF, Afzal MZ. Leveraging Grid Search in Tuning Ensemble Models to Predict Boiling Points of Alkanes. Sci Inquiry Rev [Internet]. 2025 Jun. 25 [cited 2026 Sep. 26];10(2):70-85. Available from: https://journals.umt.edu.pk/index.php/SIR/article/view/8390

Issue

Section

Mathematics

Similar Articles

1 2 3 > >> 

You may also start an advanced similarity search for this article.