PRISM: A Framework for Preprocessing Impact Analysis in Machine Learning via Signature-based Modeling
DOI:
https://doi.org/10.32350.umt-air.61.01Keywords:
Machine Learning, Data Preprocessing, Behavioral Modeling, Preprocessing Signature Tensor (PST), Preprocessing Impact Vector (PIV)Abstract
The quality and transformation of input data are crucial for Machine Learning (ML) performance, along with the model architecture. Data Preprocessing (DP) is an important factor in the learning outcome. However, it has been hardly addressed as a formally modeled transformation process. The current study argued for a new analytical framework known as PRISM (Preprocessing Impact Signature Modeling). Preprocessing can be systematically described as behavioral transformation operators that impact the process of model learning over datasets and algorithms. PRISM gives a structured interpretation of behavioral differences not only based on the differences in performance after preprocessing but also on two constructs: Preprocessing Impact Vector (PIV) and Preprocessing Signature Tensor (PST). The PIV provides performance differences in several performance measures, and the PST further extends these differences between sets of models and datasets for comparative and multi-dimensional analysis. It is evaluated on six real-world datasets, including classification and regression problems, with different ML models, such as linear models, tree-based models, ensemble models, and distance-based models. Experimental results show that there is no simple, uniform, and performance-driven preprocessing effect instead, there are structured, model-dependent, and dataset sensitive patterns. Tree-based models are insensitive to preprocessing transformations, while distance-based models are quite sensitive. Moreover, preprocessing is more consistently successful in stabilization of the model than in absolute accuracy gains, especially in the case of noise and data heterogeneity. PRISM offers an interpretable, scale-invariant, and scalable analytical framework for understanding the behavior of preprocessing in ML systems. The proposed approach moves towards a more principled and data-driven approach for preprocessing analysis and designing ML pipelines.
Downloads
References
[1] R. Barouki et al., “The COVID-19 pandemic and global environmental change: Emerging research needs,” Environ. Int., vol. 146, art. no. 106272, Jan. 2021, https://doi.org/10.1016/ j.envint.2020.106272.
[2] Z. Hammoudeh and D. Lowd, “Training data influence analysis and estimation: A survey,” Mach. Learn., vol. 113, pp. 2351–2403, March 2024, https:// doi.org/10.1007/s10994-023-06495-7.
[3] P. Koukaras and C. Tjortjis, “Data preprocessing and feature engineering for data mining: Techniques, tools, and best practices,” AI, vol. 6, no. 10, art. no. 257, Oct. 2025, https://doi.org/10.3390/ai6100257.
[4] A. A. A. Fernandes, M. Koehler, N. Konstantinou, P. Pankin, N. W. Paton, and R. Sakellariou, “Data preparation: A technological perspective and review,” SN Comput. Sci., vol. 4, art. no. 425, Jun. 2023, https://doi.org /10.1007/s42979-023-01828-8.
[5] T. Emmanuel, T. Maupong, D. Mpoeleng, T. Semong, B. Mphago, and O. Tabona, “A survey on missing data in machine learning,” J. Big Data, vol. 8, art. no. 140, Oct. 2021, https:// doi.org/10.1186/s40537-021-00516-9.
[6] W. M. Hameed and N. A. Ali, “Missing value imputation techniques: A survey,” UHD J. Sci. Technol., vol. 7, no. 1, pp. 72–81, Mar. 2023, https://doi.org/10.21928/uhdjst.v7n1y2023.pp72-81.
[7] A. D. Călin, A. M. Coroiu, and H. B. Mureşan, “Analysis of preprocessing techniques for missing data in the prediction of sunflower yield in response to the effects of climate change,” Appl. Sci., vol. 13, no. 13, art. no. 7415, Jun. 2023, https://doi.org/ 10.3390/app13137415.
[8] V. G. Biju, A.-M. Schmitt, and B. Engelmann, “Assessing the influence of sensor-induced noise on machine-learning-based changeover detection in CNC machines,” Sensors, vol. 24, no. 2, art. no. 330, Jan. 2024, https://doi.org/10.3390/s24020330.
[9] M. Usmani, Z. A. Memon, A. Zulfiqar, and R. Qureshi, “Preptimize: Automation of time series data preprocessing and forecasting,” Algorithms, vol. 17, no. 8, art. no. 332, Aug. 2024, https://doi.org/10.3390/ a17080332.
[10] C. Sancricca, G. Siracusa, and C. Cappiello, “Enhancing data preparation: Insights from a time series case study,” J. Intell. Inf. Syst., vol. 62, pp. 1503–1530, Dec. 2024, https://doi. org/10.1007/s10844-024-00867-8.
[11] K. Berahmand, F. Daneshfar, E. S. Salehi, Y. Li, and Y. Xu, “Autoencoders and their applications in machine learning: A survey,” Artif. Intell. Rev., vol. 57, no. 2, art. no. 28, 2024, https://doi.org/10.1007/s10462-023-10662-6.
[12] I. D. Lopez-Miguel, “Survey on preprocessing techniques for big data projects,” in Proc. 4th XoveTIC Conf., A Coruña, Spain, 2021, vol. 7, no. 1, art. no. 14, https://doi.org/ 10.3390/engproc2021007014.
[13] P. Kamencay, P. Hockicko, and R. Hudec, “Sensors data processing using machine learning,” Sensors, vol. 24, no. 5, art. no. 1694, Mar. 2024, https://doi.org/10.3390/s24051694.
[14] S. Dai and F. Meng, “Addressing modern and practical challenges in machine learning: A survey of online federated and transfer learning,” Appl. Intell., vol. 53, pp. 11045–11072, 2023, https://doi.org/10.1007/s10489-022-04065-3.
[15] H.-J. Park, Y.-S. Koo, H.-Y. Yang, Y.-S. Han, and C.-S. Nam, “Study on data preprocessing for machine learning based on semiconductor manufacturing processes,” Sensors, vol. 24, no. 17, art. no. 5461, Aug. 2024, https://doi.org/10.3390/s 24175461.
[16] J. N. M. Dahj and K. A. Ogudo, “Machine learning-based imputation approach with dynamic feature extraction for wireless RAN performance data preprocessing,” Symmetry, vol. 15, no. 6, art. no. 1161, May 2023, https://doi.org/10.3390 /sym15061161.
Downloads
Published
How to Cite
Issue
Section
License
Copyright (c) 2026 ABU BAKAR SHABBIR, Khadija Amber

This work is licensed under a Creative Commons Attribution 4.0 International License.
UMT-AIR follow an open-access publishing policy and full text of all published articles is available free, immediately upon publication of an issue. The journal’s contents are published and distributed under the terms of the Creative Commons Attribution 4.0 International (CC-BY 4.0) license. Thus, the work submitted to the journal implies that it is original, unpublished work of the authors (neither published previously nor accepted/under consideration for publication elsewhere). On acceptance of a manuscript for publication, a corresponding author on the behalf of all co-authors of the manuscript will sign and submit a completed the Copyright and Author Consent Form.

