Artificial Data Generation in the Implementation of Digital Twins: Techniques and Applications

Authors

  • Francisco Zenza GOVCOPP, DEGEIT, University of Aveiro, Aveiro, Portugal
  • Ana L. Ramos GOVCOPP, DEGEIT, University of Aveiro, Aveiro, Portugal
  • José V. Ferreira GOVCOPP, DEGEIT, University of Aveiro, Aveiro, Portugal
  • Luís P. Ferreira LAETA/INEGI, ISEP, Polytechnic of Porto. Dr. António Bernardino de Almeida, 431. 4249-015, Porto. Portugal
  • Ricardo Ribeiro Efacec Energia – Máquinas e Equipamentos Eléctricos, S.A., 4466-925 S. Mamede de Infesta, Portugal

DOI:

https://doi.org/10.24425/mper.2025.157218

Abstract

The growing dependence on high quality data in industrial environments drives the adoption of artificial data generation techniques, especially in the development and implementation of Digital Twins (DTs). This article presents a critical review of the main approaches to creating synthetic data, with an emphasis on their application in Industry 4.0 intelligent cyber-physical systems. Initially, traditional techniques such as Random Oversampling (ROS), SMOTE and its variants are analyzed, as well as statistical models such as the Gaussian Mixture Model (GMM). Next, Deep Learning-based methods are explored, namely Autoencoders, Variational Autoencoders and Generative Adversarial Networks (GANs), highlighting their ability to produce realistic and diverse data. The study also includes the analysis of practical cases in which DTs have been developed using synthetic data, covering domains such as wind energy, aviation and urban infrastructure. In this way, the aim of this study is to critically explore the different techniques of artificial data generation with an integration with the technology of DTs. The results suggest that the appropriate use of synthetic data can not only overcome limitations related to privacy and the scarcity of real data, but also improve the robustness and effectiveness of Digital Twins models. The article concludes by discussing the current challenges and future opportunities in integrating these techniques into smart industrial environments.

References

Aghazadeh Ardebili, A., Ficarella, A., Longo, A., Khalil, A., & Khalil, S. (2023). Hybrid Turbo-Shaft Engine Digital Twinning for Autonomous Aircraft via AI and Synthetic Data Generation. Aerospace, 10 (8), 1–17. DOI: 10.3390/aerospace10080683

Alanazi, Y., Sato, N., Ambrozewicz, P., Hiller-Blin, A., Melnitchouk, W., Battaglieri, M., Liu, T., & Li, Y. (2021). A Survey of Machine Learning-Based Physics Event Generation. IJCAI International Joint Conference on Artificial Intelligence, (Mc), 4286–4293. DOI: 10.24963/ijcai.2021/588

Almada-Lobo, F. (2015). The Industry 4.0 revolution and the future of Manufacturing Execution Systems (MES). Journal of Innovation Management, 3 (4), 16– 21. DOI: 10.24840/2183-0606_003.004_0003

Aranjuelo, N., García, S., Loyo, E., Unzueta, L., & Otaegui, O. (2021). Key strategies for synthetic data generation for training intelligent systems based on people detection from omnidirectional cameras. Computers and Electrical Engineering, 92 (July 2020), 107105. DOI: 10.1016/j.compeleceng.2021.107105

Assefa, S.A., Dervovic, D., Mahfouz, M., Tillman, R.E., Reddy, P., & Veloso, M. (2020). Generating synthetic data in finance: Opportunities, challenges and pitfalls. ICAIF 2020 – 1st ACM International Conference on AI in Finance. DOI: 10.1145/3383455.3422554

Baan, J. (2021). A Comprehensive Introduction to Bayesian Deep Learning|Towards Data Science. Retrieved June 15, 2025, from DOI: https://towards datascience.com/a-comprehensive-introduction-tobayesian-deep-learning-1221d9a051de/

Barth, R., IJsselmuiden, J.M.M., Hemming, J., & van Henten, E.J. (2017). Optimizing Realism of Synthetic Agricultural Images using Cycle Generative Adversarial Networks. Proceedings of the IEEE IROS Workshop on Agricultural Robotics, 18–22.

Batuwita, R., & Palade, V. (2010). Efficient resampling methods for training support vector machines with imbalanced datasets. Proceedings of the International Joint Conference on Neural Networks, 1–8. DOI: 10.1109/IJCNN.2010.5596787

Beregi, R., Pedone, G., Háy, B., & Váncza, J. (2021). Manufacturing execution system integration through the standardization of a common service model for cyber-physical production systems. Applied Sciences (Switzerland), 11 (16). DOI: 10.3390/app11167581

Blagus, R., & Lusa, L. (2012). Evaluation of SMOTE for high-dimensional class-imbalanced microarray data. Proceedings – 2012 11th International Conference on Machine Learning and Applications, ICMLA 2012, 2, 89–94. DOI: 10.1109/ICMLA.2012.183

Camarinha-Matos, L.M., Fornasiero, R., Ramezani, J., & Ferrada, F. (2019). Collaborative networks: A pillar of digital transformation. Applied Sciences (Switzerland), 9 (24). DOI: 10.3390/app9245431

Chawla, N.V, Bowyer, K.W., Hall, L.O., & Kegelmeyer, W.P. (2002). SMOTE : Synthetic Minority Over-sampling Technique, 16, 321–357.

Chen, J., & Little, J.J. (2019). Sports camera calibration via synthetic data. IEEE Computer Society Conference on Computer Vision and Pattern Recognition Workshops, 2019–June, 2497–2504. DOI: 10.1109/ CVPRW.2019.00305

Chryssolouris, G. (2013). Manufacturing systems: theory and practice. Springer Science & Business Media.

Cochran, D.S., Kinard, D., & Bi, Z. (2016). Manufacturing System Design Meets Big Data Analytics for Continuous Improvement. Procedia CIRP, 50, 647–652. DOI: 10.1016/j.procir.2016.05.004

Douzas, G., Bacao, F., & Last, F. (2018). Improving imbalanced learning through a heuristic oversampling method based on k-means and SMOTE. Information Sciences, 465, 1–20. DOI: 10.1016/j.ins.2018.06.056

Drummond, C., & Holte, R.C. (2003). Class Imbalance, and Cost Sensitivity: Why Under-Sampling beats Over-Sampling. Physical Review Letters, 91 (3).

El Emam, K., Mosquera, L., & Hoptroff, R. (2020). Practical synthetic data generation: balancing privacy and the broad availability of data. O’Reilly Media.

Figueira, A., & Vaz, B. (2022). Survey on Synthetic Data Generation, Evaluation Methods and GANs. Mathematics, 10 (15), 1–41. DOI: 10.3390/math10152733

Foster, D. (2022). Generative deep learning. “O’Reilly Media, Inc.

Friederich, J., Francis, D.P., Lazarova-Molnar, S., & Mohamed, N. (2022). A framework for data-driven digital twins for smart manufacturing. Computers in Industry, 136, 103586. DOI: 10.1016/j.compind.2021.103586

Gaussian mixture models. (2025). Retrieved June 15, 2025, from DOI: https://scikit-learn.org/stable/modules/mixture.html

Goodfellow, I., Bengio, Y., & Aaron, C. (2017). Deep learning. MIT Press, 521 (7553), 785. DOI: 10.1016/B978-0-12-391420-0.09987-X

Goodfellow, I.J., Pouget-abadie, J., Mirza, M., Xu, B., & Warde-farley, D. (2014). Generative Adversarial Nets, 1–9.

Han, H., Wang, W.Y., & Mao, B.H. (2005). BorderlineSMOTE: A new over-sampling method in imbalanced data sets learning. Lecture Notes in Computer Science, 3644 (PART I), 878–887. DOI: 10.1007/11538059_91

He, H., Bai, Y., Garcia, E.A., & Li, S. (2008). ADASYN: Adaptive synthetic sampling approach for imbalanced learning. Proceedings of the International Joint Conference on Neural Networks, (3), 1322–1328. DOI: 10.1109/IJCNN.2008.4633969

Jaskó, S., Skrop, A., Holczinger, T., Chován, T., & Abonyi, J. (2020). Development of manufacturing execution systems in accordance with Industry 4.0 requirements: A review of standard- and ontology-based methodologies and tools. Computers in Industry, 123. DOI: 10.1016/j.compind.2020.103300

Jo, T., & Japkowicz, N. (2004). Class imbalances versus small disjuncts. ACM SIGKDD Explorations Newsletter, 6 (1), 40–49. DOI: 10.1145/1007730.1007737

Kingma, D., & Welling, M. (2013). Auto-Encoding Variational Bayes, (Ml), 1–14.

Lan, L., You, L., Zhang, Z., Fan, Z., Zhao, W., Zeng, N., Chen, Y., & Zhou, X. (2020). Generative Adversarial Networks and Its Applications in Biomedical Informatics. Frontiers in Public Health, 8 (May), 1–14. DOI: 10.3389/fpubh.2020.00164

Leng, J., Wang, D., Shen, W., Li, X., Liu, Q., & Chen, X. (2021). Digital twins-based smart manufacturing system design in Industry 4.0: A review. Journal of Manufacturing Systems, 60 (March), 119–137. DOI: 10.1016/j.jmsy.2021.05.011

Lopes, P.V., Silveira, L., Guimaraes Aquino, R.D., Ribeiro, C.H., Skoogh, A., & Verri, F.A.N. (2024). Synthetic data generation for digital twins: enabling production systems analysis in the absence of data. International Journal of Computer Integrated Manufacturing, 37 (10–11), 1252–1269. DOI: 10.1080/ 0951192X.2024.2322981

MacQueen, J. (1967). Some methods for classification and analysis of multivariate observations. In Proceedings of the Fifth Berkeley Symposium on Mathematical Statistics and Probability, Volume 1: Statistics (Vol. 5, pp. 281–298). University of California press.

Mishra, M., Leturiondo-Zubizarreta, U., SalgadoPicón, Ó., & Galar-Pascual, D. (2015). Hybrid modelling for failure diagnosis and prognosis in the transport sector. Acquired data and synthetic data: [Modelización híbrida para el diagnóstico y pronóstico de fallos en el sector del transporte. Acquired data and synthetic data]. Dyna, 90 (2), 139–145. DOI: 10.6036/7252

Mourtzis, D. (2021). Design and operation of production networks for mass personalization in the era of cloud technology. Elsevier.

Nikolenko, S.I. (2021). Synthetic data for deep learning. Springer Optimization and Its Applications (Vol. 174). DOI: 10.1007/978-3-030-75178-4_1

Patki, N., Wedge, R., & Veeramachaneni, K. (2016). The synthetic data vault. Proceedings – 3rd IEEE International Conference on Data Science and Advanced Analytics, DSAA 2016, 399–410. DOI: 10.1109/DSAA.2016.49

Ping, H., Stoyanovich, J., & Howe, B. (2017). DataSynthesizer: Privacy-preserving synthetic datasets. ACM International Conference Proceeding Series, Part F1286. DOI: 10.1145/3085504.3091117

Pujana, A., Esteras, M., Perea, E., Maqueda, E., & Calvez, P. (2023). Hybrid-Model-Based Digital Twin of the Drivetrain of a Wind Data Generation. Energies, 16.

Qian, T., Srivastava, J., Peng, Z., & Sheu, P.C.Y. (2009). Simultaneously finding fundamental articles and new topics using a community tracking method. Lecture Notes in Computer Science (including subseries Lecture Notes in Artificial Intelligence and Lecture Notes in Bioinformatics) (Vol. 5476 LNAI). DOI: 10.1007/978-3-642-01307-2_82

Rajabi, A., & Garibay, O.O. (2022). TabFairGAN: Fair Tabular Data Generation with Generative Adversarial Networks. Machine Learning and Knowledge Extraction, 4 (2), 488–501. DOI: 10.3390/make4020022

Rios, A.J., Plevris, V., & Nogal, M. (2023). Synthetic Data Generation for the Creation of Bridge Digital Twins What-If Scenarios. COMPDYN Proceedings, 4801–4809. DOI: 10.7712/120123.10760.21262

Rubin, D.B. (1993). Statistical Disclosure Limitation (SDL). Encyclopedia of Database Systems. DOI: 10.1007/978-0-387-39940-9_3686

Russell, S.J., & Norvig, P. (2016). Artificial intelligence: a modern approach. pearson.

Schleich, B., Anwer, N., Mathieu, L., & Wartzack, S. (2017). Shaping the digital twin for design and production engineering. CIRP Annals – Manufacturing Technology, 66 (1), 141–144. DOI: 10.1016/j.cirp.2017.04.040

Segovia, M., & Garcia-Alfaro, J. (2022). Design, Modeling and Implementation of Digital Twins. Sensors, 22 (14). DOI: 10.3390/s22145396

Shao, G. (2021). Use Case Scenarios for Digital Twin Implementation Based on ISO 23247.

Shao, G., Jain, S., Laroque, C., Lee, L.H., Lendermann, P., & Rose, O. (2019). Digital Twin for Smart Manufacturing: The Simulation Aspect. Proceedings – Winter Simulation Conference, 2019-Decem (Bolton 2016), 2085–2098. DOI: 10.1109/WSC40007.2019.9004659

Tekinerdogan, B., & Verdouw, C. (2020). Systems architecture design pattern catalog for developing digital twins. Sensors (Switzerland), 20 (18), 1–20. DOI: 10.3390/s20185103

van Dinter, R., Tekinerdogan, B., & Catal, C. (2022). Predictive maintenance using digital twins: A systematic literature review. Information and Software Technology, 151 (February), 107008. DOI: 10.1016/j.infsof.2022.107008

Wright, L., & Davidson, S. (2020). How to tell the difference between a model and a digital twin. Advanced Modeling and Simulation in Engineering Sciences, 7 (1). DOI: 10.1186/s40323-020-00147-4

Xu, L., Skoularidou, M., Cuesta-Infante, A., & Veeramachaneni, K. (2019). Modeling tabular data using conditional GAN. Advances in Neural Information Processing Systems, 32 (NeurIPS).

Xu, L., & Veeramachaneni, K. (2018). Synthesizing Tabular Data using Generative Adversarial Networks. Retrieved from DOI: http://arxiv.org/abs/1811.11264

Zhan, G., Qingbo, Z., & Tingxin, S. (2014). Analysis and research on dynamic models of complex manufacturing network cascading failures. Proceedings – 2014 6th International Conference on Intelligent Human-Machine Systems and Cybernetics, IHMSC 2014, 1 (1), 388–391. DOI: 10.1109/IHMSC.2014.101

Zhang, J., Fukuda, T., & Yabuki, N. (2022). Automatic generation of synthetic datasets from a city digital twin for use in the instance segmentation of building facades. Journal of Computational Design and Engineering, 9 (5), 1737–1755. DOI: 10.1093/jcde/qwac086

Downloads

Published

2025-12-30

How to Cite

Zenza, Francisco, et al. “Artificial Data Generation in the Implementation of Digital Twins: Techniques and Applications”. Management and Production Engineering Review, vol. 16, no. 4, Dec. 2025, pp. [nr art. 9], s. 1-12, doi:10.24425/mper.2025.157218.

Issue

Section

Artykuły