Experimentation · Statistics · BanditsLG Uplus, IPTV and mobile TVTechnical write-up · October 2026 · 11 min read

Building an A/B testing platform for IPTV and mobile TV, from sample size to bandits

A walkthrough of the statistical core of an experimentation platform: how an experiment is specified, how its sample size and duration are fixed before launch, how users are randomized, how results are tested and reported, and how Bayesian testing and multi-armed bandits extend it — with simplified code for each step.

Built withPythonPower analysisRandomizationTwo-proportion z-testConfidence intervalsBayesian A/B testingMulti-armed banditsThompson sampling

IPTV and mobile TV teams change things constantly — features, screens, recommendation logic, marketing messages — and need to know which changes work. The quick method is to launch and compare the weeks after with the weeks before. That comparison cannot tell a good change from a good week: a new home screen launched in the same fortnight as a content promotion gets credit for the promotion, and one launched into a quiet week gets blamed for the quiet. Splitting traffic fixes the confounding but not everything else: a split test with no planned sample, read whenever the number first looks good, quietly inflates its own false-positive rate.

In this post we walk through the platform designed to replace that, built first for IPTV and mobile TV. It covers the life of an experiment in order: objective and metric definition, sample-size calculation from a minimum detectable effect and error rates, randomized treatment/control assignment, statistical testing with a two-proportion z-test and a confidence interval read at the planned sample, and result screens designed for interpretation by product teams rather than statisticians. Bayesian testing with Beta posteriors and multi-armed bandits with Thompson sampling were reviewed for the next stage of the platform.

Solution overview

The platform follows an experiment from plan to decision. Design happens before launch and fixes the metric, the smallest effect worth detecting and the sample per arm; assignment runs while the experiment is live; analysis happens once the planned sample is reached; and reporting presents every experiment in the same layout. The statistics sit behind the result screens, so a product team reads a lift, an interval and a decision rather than a test statistic.

Architecture
1Define

Objective and decision metric

experiment spec
2Design

Sample size and duration

power analysis
3Assign

Randomize users or sessions

randomization
4Analyze

Test and interval at the plan

z-testCI
5Report

Result screens

same layout
6Extend

Bayesian and bandit methods

Beta-BinomialThompson sampling
An experiment is defined, sized, randomized and analysed at its planned sample, then reported in a standard layout; Bayesian and bandit methods extend the analysis and allocation steps.

The numbered steps in the diagram:

  1. The team states what the change should move and the one decision metric the result will be read from.
  2. Power analysis turns the baseline, the minimum detectable effect, α and power into a sample per arm, and the eligible traffic into a test duration.
  3. Users or sessions are assigned to treatment and control at random, so promotions, weekends and drift affect both arms alike.
  4. At the planned sample, a two-proportion z-test and a confidence interval for the lift decide the result.
  5. Result screens show the lift, its interval and the test result in the same layout for every experiment.
  6. Beta posteriors give the probability that the variant wins, and Thompson sampling shifts traffic toward the leader when exposure to a worse variant is costly.

Technology stack

LayerTechnologyWhat it does here
SpecificationObjective · decision metric · randomization unit · MDE · α · powerFixes the question before any data arrive
DesignSample-size and power formulas for proportions and meansTest length known before launch
AssignmentRandom treatment/control allocationUnbiased comparison through promotions and drift
AnalysisTwo-proportion z-test · confidence interval for the liftDecision at the planned sample
ReportingResult screens for product teamsThe same layout for every experiment
ExtensionsBayesian A/B testing (Beta-Binomial) · multi-armed bandits (Thompson sampling)Reviewed for platform enhancement

Step 1: Specify the experiment before it starts

Many bad experiments are decided before any data arrive: by a vague objective, a metric chosen after the fact, or a test with no stated effect size. The platform therefore starts with a specification. The team writes down what the change is supposed to move; the one metric the decision will be read from — for a recommendation shelf, say, play-starts per session; the unit that will be randomized; the current baseline of the metric; the smallest lift worth shipping, the minimum detectable effect (MDE); and the error rates they are prepared to accept.

experiments/spec.py
from dataclasses import dataclass

@dataclass(frozen=True)
class Experiment:
    key: str                # e.g. "home-shelf-v2"
    objective: str          # what the change is supposed to move
    metric: str             # the one decision metric, e.g. play-starts per session
    metric_type: str        # "proportion" or "mean"
    unit: str               # randomization unit: "user" or "session"
    baseline: float         # current value of the metric
    mde: float              # smallest absolute lift worth shipping
    alpha: float = 0.05     # false-positive rate
    power: float = 0.80     # chance of detecting a true lift of size mde
    arms: tuple = ("A", "B")

Simplified. Field names are illustrative.

Why one decision metric? With ten independent metrics each tested at 5%, the chance that at least one looks significant by luck alone is about 40%. Other metrics can be reported; the decision is read from the one named in advance.

Step 2: Size the test with a power analysis

The sample size follows from four numbers: the baseline rate, the MDE, the false-positive rate α and the power 1 − β — the probability of detecting a true effect exactly the size of the MDE. For a conversion-type metric compared between two equal arms, the standard formula for a two-sided two-proportion test gives the sample per arm; the analogous formula with the metric's standard deviation covers means. Dividing by the eligible traffic gives the test's duration, so the team knows on launch day when the result will be read.

n per arm = ( z_(1−α/2)·√(2p̄(1−p̄)) + z_(1−β)·√(p_A(1−p_A) + p_B(1−p_B)) )² / δ²

            p_B = p_A + δ,    p̄ = (p_A + p_B) / 2,    δ = minimum detectable effect

means:      n per arm = 2 · (z_(1−α/2) + z_(1−β))² · σ² / δ²
experiments/design.py
import math
from scipy.stats import norm

def n_per_arm_proportion(p_a, mde, alpha=0.05, power=0.80):
    """Two-sided two-proportion test, equal allocation."""
    p_b = p_a + mde
    p_bar = (p_a + p_b) / 2
    z_a, z_b = norm.ppf(1 - alpha / 2), norm.ppf(power)
    root = (z_a * math.sqrt(2 * p_bar * (1 - p_bar))
            + z_b * math.sqrt(p_a * (1 - p_a) + p_b * (1 - p_b)))
    return math.ceil(root ** 2 / mde ** 2)

def n_per_arm_mean(sd, mde, alpha=0.05, power=0.80):
    z_a, z_b = norm.ppf(1 - alpha / 2), norm.ppf(power)
    return math.ceil(2 * ((z_a + z_b) * sd / mde) ** 2)

def duration_days(n_per_arm, daily_eligible, arms=2, share_in_test=1.0):
    return math.ceil(n_per_arm * arms / (daily_eligible * share_in_test))

n = n_per_arm_proportion(0.08, 0.004)     # 73,853 sessions per arm

Simplified. The example uses the live model's settings, not a real experiment.

The numbers are sobering in a useful way. With an 8% baseline, a 0.4-percentage-point MDE, α of 0.05 and 80% power — the settings in the live model below — the formula asks for about 74,000 sessions per arm. Small effects on low base rates need a lot of traffic, and knowing that before launch is what stops a team from reading noise after three days.

Step 3: Randomize assignment

Randomization is what makes the comparison fair. If every user or session is assigned to A or B at random, anything else that happens during the test — a content promotion, a holiday weekend, slow drift in the user base — affects both arms equally and cancels in the difference. For user-level experiments it also has to be stable: a user who sees the new screen on Monday and the old one on Tuesday contaminates both arms, so the same user must get the same arm every time.

experiments/assign.py
import hashlib

def assign(unit_id: str, experiment_key: str,
           arms=("A", "B"), weights=(0.5, 0.5)) -> str:
    """Random with respect to anything about the unit, identical on every call."""
    digest = hashlib.sha256(f"{experiment_key}:{unit_id}".encode()).hexdigest()
    u = int(digest[:15], 16) / 16 ** 15          # uniform in [0, 1)
    cumulative = 0.0
    for arm, w in zip(arms, weights):
        cumulative += w
        if u < cumulative:
            return arm
    return arms[-1]

Simplified. A salted hash is one common way to make assignment random and repeatable; the platform's own mechanism is not described here.

With a salted hash, different experiments use different salts, so a user's arm in one experiment says nothing about their arm in another, and no assignment table has to be looked up at serving time.

Step 4: Analyse at the planned sample

When both arms reach the planned sample, the analysis computes the lift, a two-sided two-proportion z-test and a confidence interval for the difference. The test uses the pooled proportion, as the null hypothesis of no difference implies; the interval uses each arm's own variance, because it describes the difference that was actually observed.

z  = (p̂_B − p̂_A) / √( p̄(1 − p̄)·(1/n_A + 1/n_B) ),          p̄ = pooled proportion
CI = (p̂_B − p̂_A) ± z_(1−α/2) · √( p̂_A(1−p̂_A)/n_A + p̂_B(1−p̂_B)/n_B )
experiments/analyze.py
import numpy as np
from statsmodels.stats.proportion import (confint_proportions_2indep,
                                          proportions_ztest)

def analyze(x_a, n_a, x_b, n_b, n_planned, alpha=0.05):
    """Read the experiment only once both arms reach the planned sample."""
    if min(n_a, n_b) < n_planned:
        return {"status": "running", "progress": min(n_a, n_b) / n_planned}
    z, p = proportions_ztest(np.array([x_b, x_a]), np.array([n_b, n_a]))
    lo, hi = confint_proportions_2indep(x_b, n_b, x_a, n_a, method="wald",
                                        compare="diff", alpha=alpha)
    return {"status": "decided", "lift": x_b / n_b - x_a / n_a,
            "ci": (lo, hi), "z": z, "p_value": p}

Simplified. Illustrated with statsmodels for one metric and two arms.

Why wait for the planned sample? Checking every day and stopping at the first p < 0.05 gives noise many chances to cross the line, so the real false-positive rate ends up well above the stated 5%. The running estimate can be shown; the decision is still made on schedule.

Step 5: Report results product teams can read

A test statistic is not a decision. The result screens were built for interpretation: the estimated lift with its confidence interval and p-value, presented so that a non-statistician reads the right conclusion, in the same layout for every experiment so that teams learn to read it once. Before the planned sample is reached, the experiment is shown as running rather than as a premature result.

Reading an interval against both zero and the MDE separates the cases a single p-value blurs together:

Interval for the liftWhat it saysTypical decision
Entirely above zero, estimate at or above the MDEB is better by an amount worth shippingShip B
Entirely above zero, estimate below the MDEB is better, but by less than the team said matteredJudgment call; often keep A
Includes zeroNo detectable difference at this sampleKeep A
Entirely below zeroB is worseKeep A and investigate

Step 6: Add a Bayesian read with Beta posteriors

A confidence interval is the honest summary of an experiment, and product teams often find a second number easier to act on: the probability that B is better than A. For a conversion metric this comes almost for free. Starting from a uniform Beta(1, 1) prior, each arm's conversion rate has a Beta posterior after x conversions in n sessions, and the probability that B beats A — together with an interval for the lift — is estimated by sampling from the two posteriors.

p_k | data  ~  Beta(1 + x_k, 1 + n_k − x_k)
P(p_B > p_A | data)  ≈  (1/S) · Σ_s 1[ p_B^(s) > p_A^(s) ]
experiments/bayes.py
import numpy as np

def bayes_ab(x_a, n_a, x_b, n_b, prior=(1, 1), draws=200_000, seed=0):
    rng = np.random.default_rng(seed)
    a0, b0 = prior
    p_a = rng.beta(a0 + x_a, b0 + n_a - x_a, draws)   # posterior of A's rate
    p_b = rng.beta(a0 + x_b, b0 + n_b - x_b, draws)   # posterior of B's rate
    lift = p_b - p_a
    return {"p_b_beats_a": float((lift > 0).mean()),
            "lift_mean": float(lift.mean()),
            "lift_95": tuple(np.percentile(lift, [2.5, 97.5]))}

Simplified. Monte Carlo over the two posteriors; the live model uses a normal approximation instead.

The two numbers answer different questions. The interval says how big the effect is; the posterior probability says how likely B is to be better. Both belong on the result screen, each with a sentence on what it means.

Step 7: Shift traffic with Thompson sampling when exposure is costly

A fixed 50/50 test shows the worse variant to half the users until the end. When that has a real cost — a weaker recommendation shelf during a big release, a message that annoys — a multi-armed bandit can shift traffic toward the better arm while the test runs. Thompson sampling reuses the Bayesian machinery: initialize a prior for each arm, draw a plausible conversion rate from each arm's posterior, serve the arm with the highest draw, observe the result, update the posteriors, and repeat. An arm is served in proportion to the probability that it is the best one.

experiments/bandit.py
import numpy as np

def thompson(true_rates, rounds, batch, seed=0):
    """Initialize -> sample from posteriors -> serve -> observe -> update, repeat."""
    rng = np.random.default_rng(seed)
    rates = np.asarray(true_rates)
    k = len(rates)
    a, b = np.ones(k), np.ones(k)                     # Beta(1, 1) priors
    served = np.zeros(k, dtype=int)
    for _ in range(rounds):
        theta = rng.beta(a[:, None], b[:, None], size=(k, batch))
        n = np.bincount(theta.argmax(axis=0), minlength=k)   # sessions per arm
        x = rng.binomial(n, rates)                     # simulated conversions
        a += x
        b += n - x
        served += n
    return served, a / (a + b)

served, means = thompson([0.080, 0.084], rounds=28, batch=6_000)

Simplified. A simulation of daily batches with Beta-Binomial posteriors; the rates and batch size are illustrative.

The trade-off. A bandit sends fewer users to the worse arm, at the price of an unbalanced sample, a wider final interval for the lift, and estimates that need care because allocation depended on the data. Bandits are a tool for a specific situation, not a replacement for the planned test.

Try the live model

The live model below simulates one month of sessions and reads it two ways: a launch judged before versus after, and the platform's planned, randomized test with its Bayesian and bandit read-outs.

Live simulation in your browser: one month of sessions with a weekend lift, drift and a content promotion the analyst does not know about. Left: the new screen launched to everyone on day 15 and judged before versus after with a two-proportion test. Right: the platform — a planned sample from the power formula (8% baseline, 0.4 pp minimum effect, α 0.05, 80% power), daily 50/50 randomization, the cumulative lift with its 95% band, a decision at the planned point, the probability that B beats A from a normal approximation to the Beta posteriors, and a batched Thompson-style allocation that sends each day's traffic to B in proportion to that probability, capped between 10% and 90%. Set B's true effect and draw a new month. Open the live model on its own page ↗

Results

Controlled experimentsin place of before/after comparisons for IPTV and mobile TV
Standardizeddesign, sample size, randomization and reporting across teams
ExtensibleBayesian testing and bandits reviewed for the next stage

The platform enabled IPTV and mobile TV teams to evaluate service changes using controlled experiments instead of simple before/after comparisons. It standardized experiment quality and strengthened data-driven product decisions.

Because the sample is fixed before launch, every experiment is read the same way: either it has reached its planned sample and is decided, or it has not and is still running. The Bayesian probability and the bandit extend that core for the cases that need them rather than offering a way around it.

Lessons learned

Conclusion

A before/after comparison cannot separate a change from the weeks it launched in. A platform that fixes the metric and sample before launch, randomizes assignment, decides at the planned sample with an interval and a test, and reports every experiment the same way gave IPTV and mobile TV teams a repeatable way to know whether a change worked.

Bayesian testing and Thompson sampling extend the same core rather than replace it: the posterior adds a probability teams find easy to act on, and the bandit is there for the cases where exposing users to a worse variant is the expensive part.

Limitations

About the demo and confidentiality

Sessions, conversion rates, the promotion and the effect of the variant in the embedded model are simulated. No experiment, metric definition, traffic figure, system detail or result from the real platform appears here; code is simplified and written for illustration.

Taehee Lee · Data Scientist / Technical Lead, LG Uplus (2020 – 2023)Platform design: metric definition, sample size, randomization, testing and interpretation; review of Bayesian and bandit extensions. Demo simulated for this site.