Thanks for the first open-sourced diffusion model on EHR. When we ran GAN baselines and EHRDiff on MIMIC or other datasets, we found the correlation between feature prevalence of synthetic data and feature prevalence of real data are both ~0.8, much lower than 0.99. Is there any tricks to run GAN baselines and EHRDiff?
Thanks for the first open-sourced diffusion model on EHR. When we ran GAN baselines and EHRDiff on MIMIC or other datasets, we found the correlation between feature prevalence of synthetic data and feature prevalence of real data are both ~0.8, much lower than 0.99. Is there any tricks to run GAN baselines and EHRDiff?