Traditional synthetic data generation methods rely on model-based approaches that tune the parameters of a model rather than focusing on the structure of the data itself. In contrast, Scagnostics is an exploratory graphical method that captures the structure of bivariate data using graph-theoretic measures. This paper presents a novel data generation method, scatteR, that uses Scagnostics measurements to control the characteristics of the generated dataset. By using an iterative Generalized Simulated Annealing optimizer, scatteR finds the optimal arrangement of data points that minimizes the distance between current and target Scagnostics measurements. The results demonstrate that scatteR can generate 50 data points in under 30 seconds with an average Root Mean Squared Error of 0.05, making it a useful pedagogical tool for teaching statistical methods. Overall, scatteR provides an entry point for generating datasets based on the characteristics of instance space, rather than relying on model-based simulations.
翻译:传统的合成数据生成方法依赖于基于模型的方法,即调整模型参数而非关注数据本身的结构。相比之下,Scagnostics是一种探索性图形方法,利用图论度量捕捉双变量数据的结构。本文提出了一种新颖的数据生成方法scatteR,它使用Scagnostics测量值来控制生成数据集的特性。通过使用迭代广义模拟退火优化器,scatteR找到数据点的最优排列,使当前Scagnostics测量值与目标测量值之间的距离最小化。结果表明,scatteR能在30秒内生成50个数据点,平均均方根误差为0.05,使其成为教授统计方法的有用教学工具。总体而言,scatteR提供了一种基于实例空间特征生成数据集的切入点,而非依赖基于模型的模拟。