Model Releases

Seq2Synth: Benchmarking Temporal Fidelity in Synthetic Sequential Tabular Data

arXiv:2607.15606v2 Announce Type: replace Abstract: Synthetic sequential tabular data are increasingly used for privacy-preserving data sharing and data-driven research, but evaluating their fidelity

DGX agentpaper
model-releasesarxiv-cs-lg

arXiv:2607.15606v2 Announce Type: replace Abstract: Synthetic sequential tabular data are increasingly used for privacy-preserving data sharing and data-driven research, but evaluating their fidelity remains difficult because temporal structure is easily lost under conventional tabular metrics. Existing single-table and relational evaluation protocols largely collapse records into static distributions, leaving timestamp validity, time-conditioned population structure, within-entity dynamics, and temporal relational structure insufficiently tested. We introduce Seq2Synth, a taxonomy-guided benchmark for evaluating whether synthetic sequential tabular data preserve these temporal structures. The taxonomy characterizes datasets by time representation, sampling regularity, cross-trajectory dependence, and schema structure, and determines which evaluation dimensions are applicable. Seq2Synth reports temporal fidelity across timestamp, cross-sectional, longitudinal, and structural dimensions, and extends utility and privacy evaluation to trajectory-aware settings. Across seven core datasets drawn from a 13-dataset benchmark and eight generators, static-distribution rankings diverge substantially from temporal-aware rankings; models that appear strong under conventional metrics often exhibit invalid timestamps, distorted trajectories, or inconsistent temporal relational structure. These results show that temporal fidelity cannot be inferred from static or relational metrics alone. Our benchmark code is available at https://github.com/KiwanKwon/Seq2Synth.

Source: arXiv cs.LG | 2026-08-12

Loading related sources…