Generating privacy-safe synthetic data with GANs, copulas and a utility–privacy QA report
A walkthrough of the synthetic-data work: how a table is profiled, which generator families were evaluated for tabular and time-series data and why, how a QA report decides whether a synthetic dataset is useful and safe, and how the tool was packaged and used to release certified-anonymous data — with simplified code for each step.
A telecom holds data such as IPTV viewing logs, location patterns, payment behaviour and service-usage histories. It is valuable for analysis and for data partnerships, and privacy, re-identification and legal risk keep most of it inside the company. Data for a specific purpose is also often too scarce to train on, and collecting more keeps getting more expensive. The traditional escape routes lose what made the data worth having: masking and aggregation discard detail, and shuffling each column independently keeps every distribution while destroying every relationship. A generative model can do better — and can also memorize its training rows, which defeats the purpose when the point is privacy.
In this post we walk through the synthetic-data programme the author led as technical lead. It started as research on tabular and time-series generators — Wasserstein and conditional GANs, CTGAN, CTAB-GAN and CTAB-GAN+, TabFairGAN, TabDDPM, MTcopula, TTS-GAN and TTS-CGAN — built into a web MVP on Google Cloud Run that trains a generator and produces a QA report on whether the synthetic data keeps the source's characteristics and how anonymous it is. Through the KISA data-platform project, the methods were then used to reproduce company datasets — IPTV real-time viewing logs, delivery-app usage records and origin–destination information — which were certified as anonymous and provided externally: an early practical case of synthetic data used to share data under privacy constraints.
Solution overview
The work has two paths that share one evaluation. In the MVP path, a web application on Google Cloud Run takes an uploaded table (or example data), lets the user set the variables, trains a generator, shows a QA report and offers the synthetic data for download; a time-series version takes an uploaded series and training hyperparameters, shows QA plots and generates the requested number of samples. In the release path, company datasets were reproduced with the methods developed in the research, evaluated for utility and anonymity, and certified as anonymous under the KISA data-platform program before leaving the company.
Column types and conditions
GAN, diffusion and copula generators
Transformer GANs
Sample on demand
Fidelity, utility, privacy
Web MVP
Anonymity certification
The numbered steps in the diagram:
- Each column is typed — continuous, categorical or mixed — and the user chooses the variables to reproduce or condition on.
- Tables are learned by GAN-family generators (Wasserstein GAN, conditional GAN, CTGAN, CTAB-GAN and CTAB-GAN+, TabFairGAN), a diffusion model (TabDDPM) and copula models (MTcopula).
- Sequences such as viewing logs are learned by transformer-based GANs, TTS-GAN and its conditional variant TTS-CGAN.
- Any number of rows or sequences can be drawn, optionally under conditions such as a target class.
- The QA report compares distributions and relationships with the source, trains models on synthetic data and tests them on real data, and checks how close synthetic records come to real ones.
- The loop — upload, configure, train, read the report, download — runs as a web application on Google Cloud Run.
- For external provision, reproduced datasets were evaluated for utility and anonymity and certified as anonymous through the KISA data-platform program.
Technology stack
| Layer | Technology | What it does here |
|---|---|---|
| Tabular generators | Wasserstein GAN · conditional GAN · CTGAN · CTAB-GAN · CTAB-GAN+ · TabFairGAN | Mixed continuous and categorical tables |
| Diffusion and copulas | TabDDPM · MTcopula · Gaussian copula | Alternative generators with different strengths |
| Time series | TTS-GAN · TTS-CGAN (transformer-based GANs) | Viewing logs and other sequences |
| Evaluation | Marginal and correlation fidelity · train-on-synthetic, test-on-real · nearest-record privacy checks | The evidence for release decisions |
| Application | Web MVP on Google Cloud Run | Upload → configure → train → QA report → download |
| Release | KISA data-platform program · anonymity certification | External provision of reproduced datasets |
Step 1: Profile the table before choosing a generator
A generator can only reproduce what it is told to model, so the first step fixes what each column is. Continuous columns are often skewed or multi-modal; categorical columns can be heavily imbalanced; and many telecom columns are mixed — a spike at zero (no usage this month) plus a continuous tail. Each needs a different encoding, and which ones dominate a table influences the choice of generator. The user also declares which variables must be reproduced and which, if any, generation should be conditioned on.
import pandas as pd
def profile(df: pd.DataFrame, max_levels=30, spike_share=0.05):
"""Decide how each column must be modelled before choosing a generator."""
spec = {}
for col in df.columns:
s = df[col]
if s.dtype == object or s.nunique() <= max_levels:
share = s.value_counts(normalize=True)
spec[col] = {"type": "categorical", "levels": len(share),
"rarest": float(share.min())}
else:
zero = float((s == 0).mean())
spec[col] = {"type": "mixed" if zero >= spike_share else "continuous",
"zero_share": zero, "skew": float(s.skew())}
return specSimplified. In the MVP, variable settings are confirmed by the user before training; thresholds are illustrative.
Why mixed columns matter. Treating a zero-inflated usage column as continuous smears the spike into small positive values that never occur in the source. CTAB-GAN's mixed-type encoding exists for exactly this case.
Step 2: Train tabular GANs that respect column types
A GAN trains a generator against a discriminator that tries to tell real rows from generated ones. The original objective amounts to minimizing a Jensen–Shannon divergence, which gives poor gradients when the real and generated distributions barely overlap — common early in training. The Wasserstein GAN replaces it with the earth-mover distance, the minimum amount of probability mass that must be moved to turn one distribution into the other, which stays informative even then and makes training more stable.
WGAN min_G max_(‖D‖_L ≤ 1) E_(x~real)[ D(x) ] − E_z[ D(G(z)) ] D constrained to be 1-Lipschitz
CTGAN continuous column → mode-specific normalization (variational Gaussian mixture)
discrete columns → conditional vector + training-by-sampling
Tables add problems of their own, and the tabular GANs evaluated each address one. CTGAN models every continuous column as a mixture of Gaussian modes and normalizes each value within its mode, which counters mode collapse on multi-modal columns, and conditions the generator on a randomly chosen category during training so that rare categories are learned. CTAB-GAN adds an encoding for mixed columns and a classification loss that keeps generated rows useful for a downstream target; CTAB-GAN+ adds a Wasserstein loss and further downstream losses. TabFairGAN adds a fairness measure to the generator's loss, so data can be generated with less bias against a protected attribute.
Two other families completed the comparison. TabDDPM is a diffusion model: it learns to reverse a step-by-step noising process — Gaussian noise for numerical columns, multinomial noise for categorical ones. Copula models such as MTcopula separate each column's distribution from the dependence between columns (Step 4).
from ctgan import CTGAN
def train_ctgan(train_df, spec, epochs=300):
discrete = [c for c, s in spec.items() if s["type"] == "categorical"]
model = CTGAN(epochs=epochs, batch_size=500, pac=10)
model.fit(train_df, discrete_columns=discrete)
return model
model = train_ctgan(train_df, spec)
synthetic = model.sample(len(train_df))
# conditional generation: more rows of a rare category, e.g. to rebalance a class
extra = model.sample(5_000, condition_column="segment", condition_value="rare")Simplified. Illustrated with the open-source reference implementation of CTGAN; the project's training code, settings and data are not shown.
Step 3: Generate sequences with transformer GANs
Viewing logs are sequences, and a row-by-row table generator cannot keep their order: what is watched at nine in the evening depends on what was watched at eight. For time series the work used TTS-GAN, a GAN whose generator and discriminator are both transformer encoders — the generator turns noise into a multichannel sequence, and the discriminator splits a sequence into patches and judges it the way a vision transformer judges an image — and TTS-CGAN, its conditional version, which generates sequences of a requested class.
Both learn from fixed-length windows cut from longer series and scaled to a common range. The time-series MVP took an uploaded series and the training hyperparameters, showed QA plots comparing real and generated windows, and generated the requested number of samples.
import numpy as np
def to_windows(series, length=24, stride=6):
"""(T, C) multichannel series -> (N, C, length) windows scaled to [-1, 1]."""
lo, hi = series.min(axis=0), series.max(axis=0)
span = np.where(hi > lo, hi - lo, 1.0)
scaled = 2 * (series - lo) / span - 1
starts = range(0, len(series) - length + 1, stride)
windows = np.stack([scaled[s:s + length].T for s in starts])
return windows.astype("float32"), (lo, span)
def from_windows(windows, scale):
"""Generated (N, C, length) windows back to original units, as (N, length, C)."""
lo, span = scale
return (windows.transpose(0, 2, 1) + 1) / 2 * span + loSimplified. Windowing and scaling only; the transformer generator and discriminator are not shown.
Step 4: Keep a copula as the fast baseline
Not every dataset needs a GAN. A Gaussian copula splits a table into its marginals and a dependence structure: each column is mapped to normal scores through its own empirical distribution, the correlation of those scores is estimated, and new rows are drawn from a multivariate normal and mapped back through each column's inverse distribution. Values stay in range, categories stay valid and monotone relationships between columns are kept. It fits in seconds and has almost nothing to tune.
fit z_j = Φ⁻¹( F̂_j(x_j) ), Σ̂ = corr(z) sample z ~ N(0, Σ̂), x̃_j = F̂_j⁻¹( Φ(z_j) )
What it cannot reproduce is multi-modal, conditional or non-monotone structure — the cases the GAN and diffusion families were evaluated for. That makes it a useful yardstick: a heavier generator has to beat the copula on the QA report to justify its training cost. It is also the generator in the live model on this page, standing in for the GANs used in the project.
import numpy as np
from scipy.stats import norm
def fit_copula(X, rng):
"""X: (n, d) array; categorical columns as integer codes."""
n, d = X.shape
Z = np.empty((n, d))
for j in range(d):
order = np.lexsort((rng.random(n), X[:, j])) # by value, ties at random
ranks = np.empty(n)
ranks[order] = np.arange(n)
Z[:, j] = norm.ppf((ranks + 0.5) / n) # normal scores
return {"corr": np.corrcoef(Z, rowvar=False), "sorted": np.sort(X, axis=0)}
def sample_copula(fit, m, rng):
d = fit["corr"].shape[0]
Z = rng.multivariate_normal(np.zeros(d), fit["corr"], size=m)
U = norm.cdf(Z)
n = fit["sorted"].shape[0]
idx = np.minimum((U * n).astype(int), n - 1)
return np.take_along_axis(fit["sorted"], idx, axis=0) # empirical inverse CDFSimplified. The same construction as the live model: categorical codes are sampled by quantile, so each category keeps its frequency.
Step 5: Measure fidelity and utility
The QA report answers two questions. The first is whether the synthetic data keeps the source's characteristics, checked at three levels. Marginals: each column's distribution against the source — in the sketch, a Kolmogorov–Smirnov distance for numeric columns and total variation for categorical ones. Relationships: the difference between the correlation matrices. Utility: a model trained on synthetic rows and tested on real rows that the generator never saw (train on synthetic, test on real, or TSTR), compared with the same model trained on real rows.
KS_j = sup_x | F̂_real,j(x) − F̂_syn,j(x) | numeric column j TV_j = ½ · Σ_v | P_real(x_j = v) − P_syn(x_j = v) | categorical column j TSTR = AUC( model fitted on synthetic rows, scored on a real holdout )
import numpy as np
from scipy.stats import ks_2samp
from sklearn.linear_model import LogisticRegression
from sklearn.metrics import roc_auc_score
def marginal_errors(real, syn, spec):
err = {}
for col, s in spec.items():
if s["type"] == "categorical":
p = real[col].value_counts(normalize=True)
q = syn[col].value_counts(normalize=True)
err[col] = 0.5 * p.subtract(q, fill_value=0).abs().sum() # total variation
else:
err[col] = ks_2samp(real[col], syn[col]).statistic # KS distance
return err
def correlation_error(real, syn, cols):
diff = (real[cols].corr() - syn[cols].corr()).abs().to_numpy()
upper = diff[np.triu_indices(len(cols), k=1)]
return upper.mean(), upper.max()
def tstr_auc(syn, real_train, real_holdout, features, target):
"""Train on synthetic, test on real rows the generator never saw."""
def auc(train):
clf = LogisticRegression(max_iter=1000).fit(train[features], train[target])
return roc_auc_score(real_holdout[target],
clf.predict_proba(real_holdout[features])[:, 1])
return {"trained_on_synthetic": auc(syn), "trained_on_real": auc(real_train)}Simplified. A logistic model stands in for whatever downstream model the data is meant for.
Why lead with relationships? Column shuffling passes every marginal check by construction and is useless: a model trained on shuffled rows learns nothing about real customers. An evaluation that stops at histograms will approve it, so the report leads with correlations and TSTR.
Step 6: Measure privacy against a holdout
The second question is how anonymous the data is. No exact copies of a source row is necessary and far from sufficient. The more informative test compares each synthetic row's distance to its closest real training row (DCR) with the same distance for genuinely new real rows — a holdout the generator never saw. Synthetic rows that sit no closer to training records than strangers do are not singling anyone out; a cluster of synthetic rows much closer than that is memorization, however the model was trained.
import numpy as np
from sklearn.neighbors import NearestNeighbors
from sklearn.preprocessing import StandardScaler
def privacy_report(train, holdout, syn, cols):
"""Distance to closest record (DCR): synthetic rows vs real rows never seen."""
scaler = StandardScaler().fit(train[cols])
nn = NearestNeighbors(n_neighbors=1).fit(scaler.transform(train[cols]))
d_syn = nn.kneighbors(scaler.transform(syn[cols]))[0][:, 0]
d_new = nn.kneighbors(scaler.transform(holdout[cols]))[0][:, 0]
exact = len(syn.merge(train.drop_duplicates(), how="inner"))
return {
"exact_copies": exact,
"dcr_median_synthetic": float(np.median(d_syn)),
"dcr_median_holdout": float(np.median(d_new)),
# share of synthetic rows closer than 95% of genuinely new real rows
"too_close_share": float(np.mean(d_syn < np.percentile(d_new, 5))),
}Simplified. Categorical columns are one-hot encoded before scaling; the live model uses a mixed distance instead.
These checks are evidence, not proof. For the datasets provided externally, the release decision also rested on a formal process: the reproduced data was evaluated for utility and anonymity and certified as anonymous through the KISA data-platform program before it left the company.
Step 7: Package the loop and release through review
The MVP put the whole loop behind one interface on Google Cloud Run. For tables: upload data or choose example data, set the variables, train the GAN, read the QA report on fidelity and anonymity, and download the synthetic data. For time series: upload, set the training hyperparameters, read the QA plots and generate N samples. A hosted application meant nothing to install, and the same report format applied to every dataset and every generator.
The first business use was data partnerships, where synthetic data could be applied most visibly. Company data — IPTV real-time viewing logs, delivery-app usage records and origin–destination information — was reproduced with the methods from the research, evaluated, certified as anonymous and provided externally through the KISA data-platform project, an early case in Korea of company data made shareable this way.
| Question | Evidence in the report | Fails when |
|---|---|---|
| Does each column look right? | KS / total variation per column | Ranges, spikes or rare categories are lost |
| Do the relationships survive? | Correlation difference, TSTR against real-trained | A model trained on synthetic rows is much worse on real ones |
| Is anyone reproduced? | Exact copies, nearest-record distance against a holdout | Synthetic rows sit closer to training rows than new real rows do |
Try the live model
The live model below runs the QA report in your browser on a generated subscriber table, comparing naive column shuffling with a Gaussian-copula generator — a lightweight stand-in for the GAN generators used in the project.
Live model, computed in your browser. 2,400 generated subscribers (1,800 to fit on, 600 held out) are the source. Column shuffling and a Gaussian copula — standing in for the project's GAN generators — each produce 1,800 synthetic rows, and both get the same QA report: marginal errors (KS for numeric columns, total variation for categories), correlation matrices, a logistic churn model trained on synthetic rows and tested on the real holdout, exact-copy counts, and nearest-record distances compared with the holdout's. Open the live model on its own page ↗
Results
The project converted sensitive datasets — IPTV real-time viewing logs, delivery-app usage records and origin–destination information — into externally usable synthetic data, and provided an early practical case of using synthetic data to unlock data value while addressing privacy constraints.
The web application made the method repeatable: the same profile, train, generate and QA steps applied to every dataset and generator, so whether a dataset was fit to leave the company could be read from a report rather than argued.
Lessons learned
- Utility is a relationship, not a histogram. Matching marginals is easy; the report leads with correlations and train-on-synthetic, test-on-real performance.
- Measure privacy against a holdout. Exact-copy counts miss memorization; comparing nearest-record distances with those of unseen real rows shows it.
- Keep a cheap baseline. A copula fits in seconds and sets the bar a GAN or diffusion model has to clear to justify its cost.
- One report for every model. Evaluating every generator the same way turned a list of methods into a decision about what could be released.
Conclusion
Synthetic data is only useful if it keeps the relationships analysts need and reproduces no one. Evaluating GAN, diffusion and copula generators for tables and time series against one QA report on fidelity, utility and privacy, packaging the loop as a web application, and backing release with a formal anonymity review let sensitive telecom data be provided externally as certified-anonymous synthetic data.
The same discipline — a profile, a generator chosen on evidence, and a report that tests utility and privacy against held-out data — applies to any organization that wants to share data it cannot share as it is. Planned follow-ups were location data and the use of synthetic populations in causal inference.
Limitations
- Synthetic data cannot contain more information than its source; rare segments and subtle interactions are the first to blur.
- Distance- and copy-based checks address memorization, not every possible inference attack; release decisions also rested on a formal anonymity review.
- GANs are sensitive to training settings and can be unstable; every trained model needs its own QA report rather than trust in the method.
- The live model fits a Gaussian copula on a seven-column generated table as a stand-in for the GAN generators; the project's datasets were wider, included time series, and used the GAN, diffusion and copula models described above.
About the demo and confidentiality
The source subscribers in the embedded model are themselves generated for this page from a few fixed rules. No customer data, schema, training setting, evaluation result or provided dataset from the project appears here; code is simplified and written for illustration.