Synthetic data · GANs · PrivacyLG Uplus, technical lead project · KISA data platformTechnical write-up · October 2026 · 12 min read

Generating privacy-safe synthetic data with GANs, copulas and a utility–privacy QA report

A walkthrough of the synthetic-data work: how a table is profiled, which generator families were evaluated for tabular and time-series data and why, how a QA report decides whether a synthetic dataset is useful and safe, and how the tool was packaged and used to release certified-anonymous data — with simplified code for each step.

Built withPythonWasserstein GANCTGANCTAB-GAN+TabFairGANTabDDPMTTS-GAN · TTS-CGANMTcopula · Gaussian copulaGoogle Cloud Run

A telecom holds data such as IPTV viewing logs, location patterns, payment behaviour and service-usage histories. It is valuable for analysis and for data partnerships, and privacy, re-identification and legal risk keep most of it inside the company. Data for a specific purpose is also often too scarce to train on, and collecting more keeps getting more expensive. The traditional escape routes lose what made the data worth having: masking and aggregation discard detail, and shuffling each column independently keeps every distribution while destroying every relationship. A generative model can do better — and can also memorize its training rows, which defeats the purpose when the point is privacy.

In this post we walk through the synthetic-data programme the author led as technical lead. It started as research on tabular and time-series generators — Wasserstein and conditional GANs, CTGAN, CTAB-GAN and CTAB-GAN+, TabFairGAN, TabDDPM, MTcopula, TTS-GAN and TTS-CGAN — built into a web MVP on Google Cloud Run that trains a generator and produces a QA report on whether the synthetic data keeps the source's characteristics and how anonymous it is. Through the KISA data-platform project, the methods were then used to reproduce company datasets — IPTV real-time viewing logs, delivery-app usage records and origin–destination information — which were certified as anonymous and provided externally: an early practical case of synthetic data used to share data under privacy constraints.

Solution overview

The work has two paths that share one evaluation. In the MVP path, a web application on Google Cloud Run takes an uploaded table (or example data), lets the user set the variables, trains a generator, shows a QA report and offers the synthetic data for download; a time-series version takes an uploaded series and training hyperparameters, shows QA plots and generates the requested number of samples. In the release path, company datasets were reproduced with the methods developed in the research, evaluated for utility and anonymity, and certified as anonymous under the KISA data-platform program before leaving the company.

Architecture
1Profile

Column types and conditions

Python
2Tabular

GAN, diffusion and copula generators

CTGANCTAB-GAN+TabDDPM
3Time series

Transformer GANs

TTS-GANTTS-CGAN
4Generate

Sample on demand

conditional sampling
5QA report

Fidelity, utility, privacy

TSTRnearest record
6Serve

Web MVP

Cloud Run
7Release

Anonymity certification

KISA program
A table or time series is profiled, a generator is trained and sampled, and every synthetic dataset gets the same QA report. The MVP exposes this loop as a web application; in the release path, the evidence fed a formal anonymity review before data left the company.

The numbered steps in the diagram:

  1. Each column is typed — continuous, categorical or mixed — and the user chooses the variables to reproduce or condition on.
  2. Tables are learned by GAN-family generators (Wasserstein GAN, conditional GAN, CTGAN, CTAB-GAN and CTAB-GAN+, TabFairGAN), a diffusion model (TabDDPM) and copula models (MTcopula).
  3. Sequences such as viewing logs are learned by transformer-based GANs, TTS-GAN and its conditional variant TTS-CGAN.
  4. Any number of rows or sequences can be drawn, optionally under conditions such as a target class.
  5. The QA report compares distributions and relationships with the source, trains models on synthetic data and tests them on real data, and checks how close synthetic records come to real ones.
  6. The loop — upload, configure, train, read the report, download — runs as a web application on Google Cloud Run.
  7. For external provision, reproduced datasets were evaluated for utility and anonymity and certified as anonymous through the KISA data-platform program.

Technology stack

LayerTechnologyWhat it does here
Tabular generatorsWasserstein GAN · conditional GAN · CTGAN · CTAB-GAN · CTAB-GAN+ · TabFairGANMixed continuous and categorical tables
Diffusion and copulasTabDDPM · MTcopula · Gaussian copulaAlternative generators with different strengths
Time seriesTTS-GAN · TTS-CGAN (transformer-based GANs)Viewing logs and other sequences
EvaluationMarginal and correlation fidelity · train-on-synthetic, test-on-real · nearest-record privacy checksThe evidence for release decisions
ApplicationWeb MVP on Google Cloud RunUpload → configure → train → QA report → download
ReleaseKISA data-platform program · anonymity certificationExternal provision of reproduced datasets

Step 1: Profile the table before choosing a generator

A generator can only reproduce what it is told to model, so the first step fixes what each column is. Continuous columns are often skewed or multi-modal; categorical columns can be heavily imbalanced; and many telecom columns are mixed — a spike at zero (no usage this month) plus a continuous tail. Each needs a different encoding, and which ones dominate a table influences the choice of generator. The user also declares which variables must be reproduced and which, if any, generation should be conditioned on.

synth/profile.py
import pandas as pd

def profile(df: pd.DataFrame, max_levels=30, spike_share=0.05):
    """Decide how each column must be modelled before choosing a generator."""
    spec = {}
    for col in df.columns:
        s = df[col]
        if s.dtype == object or s.nunique() <= max_levels:
            share = s.value_counts(normalize=True)
            spec[col] = {"type": "categorical", "levels": len(share),
                         "rarest": float(share.min())}
        else:
            zero = float((s == 0).mean())
            spec[col] = {"type": "mixed" if zero >= spike_share else "continuous",
                         "zero_share": zero, "skew": float(s.skew())}
    return spec

Simplified. In the MVP, variable settings are confirmed by the user before training; thresholds are illustrative.

Why mixed columns matter. Treating a zero-inflated usage column as continuous smears the spike into small positive values that never occur in the source. CTAB-GAN's mixed-type encoding exists for exactly this case.

Step 2: Train tabular GANs that respect column types

A GAN trains a generator against a discriminator that tries to tell real rows from generated ones. The original objective amounts to minimizing a Jensen–Shannon divergence, which gives poor gradients when the real and generated distributions barely overlap — common early in training. The Wasserstein GAN replaces it with the earth-mover distance, the minimum amount of probability mass that must be moved to turn one distribution into the other, which stays informative even then and makes training more stable.

WGAN    min_G  max_(‖D‖_L ≤ 1)   E_(x~real)[ D(x) ] − E_z[ D(G(z)) ]        D constrained to be 1-Lipschitz

CTGAN   continuous column → mode-specific normalization (variational Gaussian mixture)
        discrete columns  → conditional vector + training-by-sampling

Tables add problems of their own, and the tabular GANs evaluated each address one. CTGAN models every continuous column as a mixture of Gaussian modes and normalizes each value within its mode, which counters mode collapse on multi-modal columns, and conditions the generator on a randomly chosen category during training so that rare categories are learned. CTAB-GAN adds an encoding for mixed columns and a classification loss that keeps generated rows useful for a downstream target; CTAB-GAN+ adds a Wasserstein loss and further downstream losses. TabFairGAN adds a fairness measure to the generator's loss, so data can be generated with less bias against a protected attribute.

Two other families completed the comparison. TabDDPM is a diffusion model: it learns to reverse a step-by-step noising process — Gaussian noise for numerical columns, multinomial noise for categorical ones. Copula models such as MTcopula separate each column's distribution from the dependence between columns (Step 4).

synth/tabular.py
from ctgan import CTGAN

def train_ctgan(train_df, spec, epochs=300):
    discrete = [c for c, s in spec.items() if s["type"] == "categorical"]
    model = CTGAN(epochs=epochs, batch_size=500, pac=10)
    model.fit(train_df, discrete_columns=discrete)
    return model

model = train_ctgan(train_df, spec)
synthetic = model.sample(len(train_df))

# conditional generation: more rows of a rare category, e.g. to rebalance a class
extra = model.sample(5_000, condition_column="segment", condition_value="rare")

Simplified. Illustrated with the open-source reference implementation of CTGAN; the project's training code, settings and data are not shown.

Step 3: Generate sequences with transformer GANs

Viewing logs are sequences, and a row-by-row table generator cannot keep their order: what is watched at nine in the evening depends on what was watched at eight. For time series the work used TTS-GAN, a GAN whose generator and discriminator are both transformer encoders — the generator turns noise into a multichannel sequence, and the discriminator splits a sequence into patches and judges it the way a vision transformer judges an image — and TTS-CGAN, its conditional version, which generates sequences of a requested class.

Both learn from fixed-length windows cut from longer series and scaled to a common range. The time-series MVP took an uploaded series and the training hyperparameters, showed QA plots comparing real and generated windows, and generated the requested number of samples.

synth/sequences.py
import numpy as np

def to_windows(series, length=24, stride=6):
    """(T, C) multichannel series -> (N, C, length) windows scaled to [-1, 1]."""
    lo, hi = series.min(axis=0), series.max(axis=0)
    span = np.where(hi > lo, hi - lo, 1.0)
    scaled = 2 * (series - lo) / span - 1
    starts = range(0, len(series) - length + 1, stride)
    windows = np.stack([scaled[s:s + length].T for s in starts])
    return windows.astype("float32"), (lo, span)

def from_windows(windows, scale):
    """Generated (N, C, length) windows back to original units, as (N, length, C)."""
    lo, span = scale
    return (windows.transpose(0, 2, 1) + 1) / 2 * span + lo

Simplified. Windowing and scaling only; the transformer generator and discriminator are not shown.

Step 4: Keep a copula as the fast baseline

Not every dataset needs a GAN. A Gaussian copula splits a table into its marginals and a dependence structure: each column is mapped to normal scores through its own empirical distribution, the correlation of those scores is estimated, and new rows are drawn from a multivariate normal and mapped back through each column's inverse distribution. Values stay in range, categories stay valid and monotone relationships between columns are kept. It fits in seconds and has almost nothing to tune.

fit      z_j = Φ⁻¹( F̂_j(x_j) ),      Σ̂ = corr(z)
sample   z ~ N(0, Σ̂),      x̃_j = F̂_j⁻¹( Φ(z_j) )

What it cannot reproduce is multi-modal, conditional or non-monotone structure — the cases the GAN and diffusion families were evaluated for. That makes it a useful yardstick: a heavier generator has to beat the copula on the QA report to justify its training cost. It is also the generator in the live model on this page, standing in for the GANs used in the project.

synth/copula.py
import numpy as np
from scipy.stats import norm

def fit_copula(X, rng):
    """X: (n, d) array; categorical columns as integer codes."""
    n, d = X.shape
    Z = np.empty((n, d))
    for j in range(d):
        order = np.lexsort((rng.random(n), X[:, j]))   # by value, ties at random
        ranks = np.empty(n)
        ranks[order] = np.arange(n)
        Z[:, j] = norm.ppf((ranks + 0.5) / n)           # normal scores
    return {"corr": np.corrcoef(Z, rowvar=False), "sorted": np.sort(X, axis=0)}

def sample_copula(fit, m, rng):
    d = fit["corr"].shape[0]
    Z = rng.multivariate_normal(np.zeros(d), fit["corr"], size=m)
    U = norm.cdf(Z)
    n = fit["sorted"].shape[0]
    idx = np.minimum((U * n).astype(int), n - 1)
    return np.take_along_axis(fit["sorted"], idx, axis=0)   # empirical inverse CDF

Simplified. The same construction as the live model: categorical codes are sampled by quantile, so each category keeps its frequency.

Step 5: Measure fidelity and utility

The QA report answers two questions. The first is whether the synthetic data keeps the source's characteristics, checked at three levels. Marginals: each column's distribution against the source — in the sketch, a Kolmogorov–Smirnov distance for numeric columns and total variation for categorical ones. Relationships: the difference between the correlation matrices. Utility: a model trained on synthetic rows and tested on real rows that the generator never saw (train on synthetic, test on real, or TSTR), compared with the same model trained on real rows.

KS_j   = sup_x | F̂_real,j(x) − F̂_syn,j(x) |            numeric column j
TV_j   = ½ · Σ_v | P_real(x_j = v) − P_syn(x_j = v) |      categorical column j
TSTR   = AUC( model fitted on synthetic rows, scored on a real holdout )
qa/fidelity.py
import numpy as np
from scipy.stats import ks_2samp
from sklearn.linear_model import LogisticRegression
from sklearn.metrics import roc_auc_score

def marginal_errors(real, syn, spec):
    err = {}
    for col, s in spec.items():
        if s["type"] == "categorical":
            p = real[col].value_counts(normalize=True)
            q = syn[col].value_counts(normalize=True)
            err[col] = 0.5 * p.subtract(q, fill_value=0).abs().sum()  # total variation
        else:
            err[col] = ks_2samp(real[col], syn[col]).statistic         # KS distance
    return err

def correlation_error(real, syn, cols):
    diff = (real[cols].corr() - syn[cols].corr()).abs().to_numpy()
    upper = diff[np.triu_indices(len(cols), k=1)]
    return upper.mean(), upper.max()

def tstr_auc(syn, real_train, real_holdout, features, target):
    """Train on synthetic, test on real rows the generator never saw."""
    def auc(train):
        clf = LogisticRegression(max_iter=1000).fit(train[features], train[target])
        return roc_auc_score(real_holdout[target],
                             clf.predict_proba(real_holdout[features])[:, 1])
    return {"trained_on_synthetic": auc(syn), "trained_on_real": auc(real_train)}

Simplified. A logistic model stands in for whatever downstream model the data is meant for.

Why lead with relationships? Column shuffling passes every marginal check by construction and is useless: a model trained on shuffled rows learns nothing about real customers. An evaluation that stops at histograms will approve it, so the report leads with correlations and TSTR.

Step 6: Measure privacy against a holdout

The second question is how anonymous the data is. No exact copies of a source row is necessary and far from sufficient. The more informative test compares each synthetic row's distance to its closest real training row (DCR) with the same distance for genuinely new real rows — a holdout the generator never saw. Synthetic rows that sit no closer to training records than strangers do are not singling anyone out; a cluster of synthetic rows much closer than that is memorization, however the model was trained.

qa/privacy.py
import numpy as np
from sklearn.neighbors import NearestNeighbors
from sklearn.preprocessing import StandardScaler

def privacy_report(train, holdout, syn, cols):
    """Distance to closest record (DCR): synthetic rows vs real rows never seen."""
    scaler = StandardScaler().fit(train[cols])
    nn = NearestNeighbors(n_neighbors=1).fit(scaler.transform(train[cols]))
    d_syn = nn.kneighbors(scaler.transform(syn[cols]))[0][:, 0]
    d_new = nn.kneighbors(scaler.transform(holdout[cols]))[0][:, 0]
    exact = len(syn.merge(train.drop_duplicates(), how="inner"))
    return {
        "exact_copies": exact,
        "dcr_median_synthetic": float(np.median(d_syn)),
        "dcr_median_holdout": float(np.median(d_new)),
        # share of synthetic rows closer than 95% of genuinely new real rows
        "too_close_share": float(np.mean(d_syn < np.percentile(d_new, 5))),
    }

Simplified. Categorical columns are one-hot encoded before scaling; the live model uses a mixed distance instead.

These checks are evidence, not proof. For the datasets provided externally, the release decision also rested on a formal process: the reproduced data was evaluated for utility and anonymity and certified as anonymous through the KISA data-platform program before it left the company.

Step 7: Package the loop and release through review

The MVP put the whole loop behind one interface on Google Cloud Run. For tables: upload data or choose example data, set the variables, train the GAN, read the QA report on fidelity and anonymity, and download the synthetic data. For time series: upload, set the training hyperparameters, read the QA plots and generate N samples. A hosted application meant nothing to install, and the same report format applied to every dataset and every generator.

The first business use was data partnerships, where synthetic data could be applied most visibly. Company data — IPTV real-time viewing logs, delivery-app usage records and origin–destination information — was reproduced with the methods from the research, evaluated, certified as anonymous and provided externally through the KISA data-platform project, an early case in Korea of company data made shareable this way.

QuestionEvidence in the reportFails when
Does each column look right?KS / total variation per columnRanges, spikes or rare categories are lost
Do the relationships survive?Correlation difference, TSTR against real-trainedA model trained on synthetic rows is much worse on real ones
Is anyone reproduced?Exact copies, nearest-record distance against a holdoutSynthetic rows sit closer to training rows than new real rows do

Try the live model

The live model below runs the QA report in your browser on a generated subscriber table, comparing naive column shuffling with a Gaussian-copula generator — a lightweight stand-in for the GAN generators used in the project.

Live model, computed in your browser. 2,400 generated subscribers (1,800 to fit on, 600 held out) are the source. Column shuffling and a Gaussian copula — standing in for the project's GAN generators — each produce 1,800 synthetic rows, and both get the same QA report: marginal errors (KS for numeric columns, total variation for categories), correlation matrices, a logistic churn model trained on synthetic rows and tested on the real holdout, exact-copy counts, and nearest-record distances compared with the holdout's. Open the live model on its own page ↗

Results

Certifiedanonymous reproduced data provided through the KISA data-platform program
Threesensitive sources made shareable: IPTV real-time viewing, delivery-app usage, origin–destination data
Tengenerative methods evaluated for tabular and time-series data

The project converted sensitive datasets — IPTV real-time viewing logs, delivery-app usage records and origin–destination information — into externally usable synthetic data, and provided an early practical case of using synthetic data to unlock data value while addressing privacy constraints.

The web application made the method repeatable: the same profile, train, generate and QA steps applied to every dataset and generator, so whether a dataset was fit to leave the company could be read from a report rather than argued.

Lessons learned

Conclusion

Synthetic data is only useful if it keeps the relationships analysts need and reproduces no one. Evaluating GAN, diffusion and copula generators for tables and time series against one QA report on fidelity, utility and privacy, packaging the loop as a web application, and backing release with a formal anonymity review let sensitive telecom data be provided externally as certified-anonymous synthetic data.

The same discipline — a profile, a generator chosen on evidence, and a report that tests utility and privacy against held-out data — applies to any organization that wants to share data it cannot share as it is. Planned follow-ups were location data and the use of synthetic populations in causal inference.

Limitations

About the demo and confidentiality

The source subscribers in the embedded model are themselves generated for this page from a few fixed rules. No customer data, schema, training setting, evaluation result or provided dataset from the project appears here; code is simplified and written for illustration.

Taehee Lee · Data Scientist / Technical Lead, LG Uplus (2020 – 2023)Technical lead: method research, prototyping, evaluation framework and the KISA data-provision work. Demo re-implemented on generated data for this site.