Artificial Data Generation in the Implementation of Digital Twins: Techniques and Applications
DOI:
https://doi.org/10.24425/mper.2025.157218Abstract
The growing dependence on high quality data in industrial environments drives the adoption of artificial data generation techniques, especially in the development and implementation of Digital Twins (DTs). This article presents a critical review of the main approaches to creating synthetic data, with an emphasis on their application in Industry 4.0 intelligent cyber-physical systems. Initially, traditional techniques such as Random Oversampling (ROS), SMOTE and its variants are analyzed, as well as statistical models such as the Gaussian Mixture Model (GMM). Next, Deep Learning-based methods are explored, namely Autoencoders, Variational Autoencoders and Generative Adversarial Networks (GANs), highlighting their ability to produce realistic and diverse data. The study also includes the analysis of practical cases in which DTs have been developed using synthetic data, covering domains such as wind energy, aviation and urban infrastructure. In this way, the aim of this study is to critically explore the different techniques of artificial data generation with an integration with the technology of DTs. The results suggest that the appropriate use of synthetic data can not only overcome limitations related to privacy and the scarcity of real data, but also improve the robustness and effectiveness of Digital Twins models. The article concludes by discussing the current challenges and future opportunities in integrating these techniques into smart industrial environments.References
Aghazadeh Ardebili, A., Ficarella, A., Longo, A., Khalil, A., & Khalil, S. (2023). Hybrid Turbo-Shaft Engine Digital Twinning for Autonomous Aircraft via AI and Synthetic Data Generation. Aerospace, 10 (8), 1–17. DOI: 10.3390/aerospace10080683
Alanazi, Y., Sato, N., Ambrozewicz, P., Hiller-Blin, A., Melnitchouk, W., Battaglieri, M., Liu, T., & Li, Y. (2021). A Survey of Machine Learning-Based Physics Event Generation. IJCAI International Joint Conference on Artificial Intelligence, (Mc), 4286–4293. DOI: 10.24963/ijcai.2021/588
Almada-Lobo, F. (2015). The Industry 4.0 revolution and the future of Manufacturing Execution Systems (MES). Journal of Innovation Management, 3 (4), 16– 21. DOI: 10.24840/2183-0606_003.004_0003
Aranjuelo, N., García, S., Loyo, E., Unzueta, L., & Otaegui, O. (2021). Key strategies for synthetic data generation for training intelligent systems based on people detection from omnidirectional cameras. Computers and Electrical Engineering, 92 (July 2020), 107105. DOI: 10.1016/j.compeleceng.2021.107105
Assefa, S.A., Dervovic, D., Mahfouz, M., Tillman, R.E., Reddy, P., & Veloso, M. (2020). Generating synthetic data in finance: Opportunities, challenges and pitfalls. ICAIF 2020 – 1st ACM International Conference on AI in Finance. DOI: 10.1145/3383455.3422554
Baan, J. (2021). A Comprehensive Introduction to Bayesian Deep Learning|Towards Data Science. Retrieved June 15, 2025, from DOI: https://towards datascience.com/a-comprehensive-introduction-tobayesian-deep-learning-1221d9a051de/
Barth, R., IJsselmuiden, J.M.M., Hemming, J., & van Henten, E.J. (2017). Optimizing Realism of Synthetic Agricultural Images using Cycle Generative Adversarial Networks. Proceedings of the IEEE IROS Workshop on Agricultural Robotics, 18–22.
Batuwita, R., & Palade, V. (2010). Efficient resampling methods for training support vector machines with imbalanced datasets. Proceedings of the International Joint Conference on Neural Networks, 1–8. DOI: 10.1109/IJCNN.2010.5596787
Beregi, R., Pedone, G., Háy, B., & Váncza, J. (2021). Manufacturing execution system integration through the standardization of a common service model for cyber-physical production systems. Applied Sciences (Switzerland), 11 (16). DOI: 10.3390/app11167581
Blagus, R., & Lusa, L. (2012). Evaluation of SMOTE for high-dimensional class-imbalanced microarray data. Proceedings – 2012 11th International Conference on Machine Learning and Applications, ICMLA 2012, 2, 89–94. DOI: 10.1109/ICMLA.2012.183
Camarinha-Matos, L.M., Fornasiero, R., Ramezani, J., & Ferrada, F. (2019). Collaborative networks: A pillar of digital transformation. Applied Sciences (Switzerland), 9 (24). DOI: 10.3390/app9245431
Chawla, N.V, Bowyer, K.W., Hall, L.O., & Kegelmeyer, W.P. (2002). SMOTE : Synthetic Minority Over-sampling Technique, 16, 321–357.
Chen, J., & Little, J.J. (2019). Sports camera calibration via synthetic data. IEEE Computer Society Conference on Computer Vision and Pattern Recognition Workshops, 2019–June, 2497–2504. DOI: 10.1109/ CVPRW.2019.00305
Chryssolouris, G. (2013). Manufacturing systems: theory and practice. Springer Science & Business Media.
Cochran, D.S., Kinard, D., & Bi, Z. (2016). Manufacturing System Design Meets Big Data Analytics for Continuous Improvement. Procedia CIRP, 50, 647–652. DOI: 10.1016/j.procir.2016.05.004
Douzas, G., Bacao, F., & Last, F. (2018). Improving imbalanced learning through a heuristic oversampling method based on k-means and SMOTE. Information Sciences, 465, 1–20. DOI: 10.1016/j.ins.2018.06.056
Drummond, C., & Holte, R.C. (2003). Class Imbalance, and Cost Sensitivity: Why Under-Sampling beats Over-Sampling. Physical Review Letters, 91 (3).
El Emam, K., Mosquera, L., & Hoptroff, R. (2020). Practical synthetic data generation: balancing privacy and the broad availability of data. O’Reilly Media.
Figueira, A., & Vaz, B. (2022). Survey on Synthetic Data Generation, Evaluation Methods and GANs. Mathematics, 10 (15), 1–41. DOI: 10.3390/math10152733
Foster, D. (2022). Generative deep learning. “O’Reilly Media, Inc.
Friederich, J., Francis, D.P., Lazarova-Molnar, S., & Mohamed, N. (2022). A framework for data-driven digital twins for smart manufacturing. Computers in Industry, 136, 103586. DOI: 10.1016/j.compind.2021.103586
Gaussian mixture models. (2025). Retrieved June 15, 2025, from DOI: https://scikit-learn.org/stable/modules/mixture.html
Goodfellow, I., Bengio, Y., & Aaron, C. (2017). Deep learning. MIT Press, 521 (7553), 785. DOI: 10.1016/B978-0-12-391420-0.09987-X
Goodfellow, I.J., Pouget-abadie, J., Mirza, M., Xu, B., & Warde-farley, D. (2014). Generative Adversarial Nets, 1–9.
Han, H., Wang, W.Y., & Mao, B.H. (2005). BorderlineSMOTE: A new over-sampling method in imbalanced data sets learning. Lecture Notes in Computer Science, 3644 (PART I), 878–887. DOI: 10.1007/11538059_91
He, H., Bai, Y., Garcia, E.A., & Li, S. (2008). ADASYN: Adaptive synthetic sampling approach for imbalanced learning. Proceedings of the International Joint Conference on Neural Networks, (3), 1322–1328. DOI: 10.1109/IJCNN.2008.4633969
Jaskó, S., Skrop, A., Holczinger, T., Chován, T., & Abonyi, J. (2020). Development of manufacturing execution systems in accordance with Industry 4.0 requirements: A review of standard- and ontology-based methodologies and tools. Computers in Industry, 123. DOI: 10.1016/j.compind.2020.103300
Jo, T., & Japkowicz, N. (2004). Class imbalances versus small disjuncts. ACM SIGKDD Explorations Newsletter, 6 (1), 40–49. DOI: 10.1145/1007730.1007737
Kingma, D., & Welling, M. (2013). Auto-Encoding Variational Bayes, (Ml), 1–14.
Lan, L., You, L., Zhang, Z., Fan, Z., Zhao, W., Zeng, N., Chen, Y., & Zhou, X. (2020). Generative Adversarial Networks and Its Applications in Biomedical Informatics. Frontiers in Public Health, 8 (May), 1–14. DOI: 10.3389/fpubh.2020.00164
Leng, J., Wang, D., Shen, W., Li, X., Liu, Q., & Chen, X. (2021). Digital twins-based smart manufacturing system design in Industry 4.0: A review. Journal of Manufacturing Systems, 60 (March), 119–137. DOI: 10.1016/j.jmsy.2021.05.011
Lopes, P.V., Silveira, L., Guimaraes Aquino, R.D., Ribeiro, C.H., Skoogh, A., & Verri, F.A.N. (2024). Synthetic data generation for digital twins: enabling production systems analysis in the absence of data. International Journal of Computer Integrated Manufacturing, 37 (10–11), 1252–1269. DOI: 10.1080/ 0951192X.2024.2322981
MacQueen, J. (1967). Some methods for classification and analysis of multivariate observations. In Proceedings of the Fifth Berkeley Symposium on Mathematical Statistics and Probability, Volume 1: Statistics (Vol. 5, pp. 281–298). University of California press.
Mishra, M., Leturiondo-Zubizarreta, U., SalgadoPicón, Ó., & Galar-Pascual, D. (2015). Hybrid modelling for failure diagnosis and prognosis in the transport sector. Acquired data and synthetic data: [Modelización híbrida para el diagnóstico y pronóstico de fallos en el sector del transporte. Acquired data and synthetic data]. Dyna, 90 (2), 139–145. DOI: 10.6036/7252
Mourtzis, D. (2021). Design and operation of production networks for mass personalization in the era of cloud technology. Elsevier.
Nikolenko, S.I. (2021). Synthetic data for deep learning. Springer Optimization and Its Applications (Vol. 174). DOI: 10.1007/978-3-030-75178-4_1
Patki, N., Wedge, R., & Veeramachaneni, K. (2016). The synthetic data vault. Proceedings – 3rd IEEE International Conference on Data Science and Advanced Analytics, DSAA 2016, 399–410. DOI: 10.1109/DSAA.2016.49
Ping, H., Stoyanovich, J., & Howe, B. (2017). DataSynthesizer: Privacy-preserving synthetic datasets. ACM International Conference Proceeding Series, Part F1286. DOI: 10.1145/3085504.3091117
Pujana, A., Esteras, M., Perea, E., Maqueda, E., & Calvez, P. (2023). Hybrid-Model-Based Digital Twin of the Drivetrain of a Wind Data Generation. Energies, 16.
Qian, T., Srivastava, J., Peng, Z., & Sheu, P.C.Y. (2009). Simultaneously finding fundamental articles and new topics using a community tracking method. Lecture Notes in Computer Science (including subseries Lecture Notes in Artificial Intelligence and Lecture Notes in Bioinformatics) (Vol. 5476 LNAI). DOI: 10.1007/978-3-642-01307-2_82
Rajabi, A., & Garibay, O.O. (2022). TabFairGAN: Fair Tabular Data Generation with Generative Adversarial Networks. Machine Learning and Knowledge Extraction, 4 (2), 488–501. DOI: 10.3390/make4020022
Rios, A.J., Plevris, V., & Nogal, M. (2023). Synthetic Data Generation for the Creation of Bridge Digital Twins What-If Scenarios. COMPDYN Proceedings, 4801–4809. DOI: 10.7712/120123.10760.21262
Rubin, D.B. (1993). Statistical Disclosure Limitation (SDL). Encyclopedia of Database Systems. DOI: 10.1007/978-0-387-39940-9_3686
Russell, S.J., & Norvig, P. (2016). Artificial intelligence: a modern approach. pearson.
Schleich, B., Anwer, N., Mathieu, L., & Wartzack, S. (2017). Shaping the digital twin for design and production engineering. CIRP Annals – Manufacturing Technology, 66 (1), 141–144. DOI: 10.1016/j.cirp.2017.04.040
Segovia, M., & Garcia-Alfaro, J. (2022). Design, Modeling and Implementation of Digital Twins. Sensors, 22 (14). DOI: 10.3390/s22145396
Shao, G. (2021). Use Case Scenarios for Digital Twin Implementation Based on ISO 23247.
Shao, G., Jain, S., Laroque, C., Lee, L.H., Lendermann, P., & Rose, O. (2019). Digital Twin for Smart Manufacturing: The Simulation Aspect. Proceedings – Winter Simulation Conference, 2019-Decem (Bolton 2016), 2085–2098. DOI: 10.1109/WSC40007.2019.9004659
Tekinerdogan, B., & Verdouw, C. (2020). Systems architecture design pattern catalog for developing digital twins. Sensors (Switzerland), 20 (18), 1–20. DOI: 10.3390/s20185103
van Dinter, R., Tekinerdogan, B., & Catal, C. (2022). Predictive maintenance using digital twins: A systematic literature review. Information and Software Technology, 151 (February), 107008. DOI: 10.1016/j.infsof.2022.107008
Wright, L., & Davidson, S. (2020). How to tell the difference between a model and a digital twin. Advanced Modeling and Simulation in Engineering Sciences, 7 (1). DOI: 10.1186/s40323-020-00147-4
Xu, L., Skoularidou, M., Cuesta-Infante, A., & Veeramachaneni, K. (2019). Modeling tabular data using conditional GAN. Advances in Neural Information Processing Systems, 32 (NeurIPS).
Xu, L., & Veeramachaneni, K. (2018). Synthesizing Tabular Data using Generative Adversarial Networks. Retrieved from DOI: http://arxiv.org/abs/1811.11264
Zhan, G., Qingbo, Z., & Tingxin, S. (2014). Analysis and research on dynamic models of complex manufacturing network cascading failures. Proceedings – 2014 6th International Conference on Intelligent Human-Machine Systems and Cybernetics, IHMSC 2014, 1 (1), 388–391. DOI: 10.1109/IHMSC.2014.101
Zhang, J., Fukuda, T., & Yabuki, N. (2022). Automatic generation of synthetic datasets from a city digital twin for use in the instance segmentation of building facades. Journal of Computational Design and Engineering, 9 (5), 1737–1755. DOI: 10.1093/jcde/qwac086