UDC 330.4
DOI: 10.36871/ek.up.p.r.2025.06.08.020

Authors

Kamil I. Shigapov,
Kazan National Research Technical University named after A. N. Tupolev, Kazan, Russia
Zarina M. Lorsanova,
A. A. Kadyrov Chechen State University, Grozny, Russia
Anna A. Ilmushkina,
Russian Biotechnological University, Moscow, Russia

Abstract

The article discusses modern methods of synthetic data supply used to train machine learning models in conditions of sample shortage, typical for economic research and practice. The main approaches to generating artificial data are analyzed, including methods based on statistical models, variational autoencoders and generative adversarial networks. Particular attention is paid to assessing the quality of synthetic data and their impact on the accuracy and stability of economic models. The need to use synthetic data to increase the representativeness of training samples, reduce the risk of overfitting and ensure the confidentiality of the original data is substantiated. Recommendations are proposed for integrating synthetic data into economic analytical processes, taking into account the specifics of economic information and the requirements for its reliability.

Keywords

synthetic data, machine learning, sample shortage, economic modeling, data generation, data quality