Forecasting · XGBoost · ProphetLG Uplus, IPTV movie businessTechnical write-up · October 2026 · 12 min read

Forecasting IPTV movie revenue with XGBoost, curve clustering, Bass diffusion and Prophet

A walkthrough of the three forecasts behind IPTV movie decisions — what a title will earn, how those sales will unfold after release, and where the whole market is heading — covering the data set, the models, how they were evaluated and why each was chosen, with simplified code for each step.

Built withPythonpandasscikit-learnXGBoostLightGBMRandom forestk-means · DTWBass diffusionProphetKOBIS public data

IPTV movie decisions — which titles to license, which to offer free, when to promote and how to split the marketing budget — all rest on expected revenue, and they ask three different questions. Before a licence is signed: how much will this title earn? After release: how will that total unfold week by week, so that promotion is timed to the curve rather than against it? Across the catalogue: where is the whole market heading? A rule of thumb such as scaling the catalogue's average curve answers none of them well. VOD sales are front-loaded and volatile, a blockbuster and a word-of-mouth title can look alike for their first week, and the market total moves with weekdays, holidays and the release slate as much as with any single title.

In this post we walk through the forecasting system built instead. It splits the problem into models that share one title-level data set: an XGBoost regressor tuned by grid search and five-fold cross-validation for a title's total; curve-type clustering of eight-week sales curves (Euclidean and DTW k-means compared) with a multiclass XGBoost that predicts a new title's curve type, plus the Bass diffusion model fitted to early sales for its trajectory; and Prophet with Korean holidays and a release-event regressor for the market's daily total. Every model is scored with volume-weighted errors over rolling forecast origins, because the titles that sell most are where acquisition and promotion money goes.

Solution overview

The system is a set of offline Python models over one data set: VOD purchase history joined with title metadata and public theatrical data. The two pre-release models — total and curve type — are trained on the back catalogue and applied to titles under consideration or just released; the trajectory model is refitted as a title's sales arrive; the market model is refitted on the daily total and forecasts several weeks ahead. Each model answers one question and is evaluated on its own terms.

Architecture
1Data

One title-level data set

pandasKOBIS data
2Title total

Revenue over the selling window

XGBoostLightGBMRandom forest
3Curve type

Cluster eight-week sales curves

k-meansDTW
4Classify

Predict a new title's curve type

XGBoost multiclass
5Trajectory

Bass curve from early sales

Bass modelleast squares
6Market

Daily total with known events

Prophetholidays
7Evaluate

Weighted errors, rolling origins

wMAPEmBRSE
One title-level data set feeds two pre-release models (total and curve type) and a trajectory model refitted as sales arrive; the market model runs on the daily total. Every model is scored with volume-weighted errors over many forecast origins.

The numbered steps in the diagram:

  1. VOD purchase history is aggregated per title with pandas and joined with title metadata, theatrical box office and admissions from the public KOBIS data, and audience ratings.
  2. XGBoost predicts each title's revenue (on a log scale) from pre-release features, tuned by grid search with five-fold cross-validation and compared with random forest and LightGBM; permutation importance names the drivers.
  3. Each title's first eight weeks become a sales-ratio curve, clustered with k-means under Euclidean distance and dynamic time warping; the number of clusters is chosen from within-cluster sum of squares, silhouette and Davies–Bouldin scores.
  4. A multiclass XGBoost with a cross-entropy loss predicts which curve type a new title will follow, from the same pre-release features.
  5. Once a title is on sale, the Bass diffusion model — a market size plus innovation and imitation coefficients — is fitted to its early sales by nonlinear least squares.
  6. Prophet decomposes total daily sales into trend, weekly and yearly seasonality, Korean holidays and a regressor for major releases; the market forecast is finished as an ensemble.
  7. Models are compared with RMSE, weighted MAPE, mBRSE and weighted accuracy, recomputed over many rolling forecast origins.

Technology stack

LayerTechnologyWhat it does here
DataPython · pandas · public box-office data (KOBIS) · audience ratingsTitle-level features joined to VOD sales
Title totalXGBoost (grid search, 5-fold CV) · random forest · LightGBM · permutation importanceRevenue over the selling window, and what drives it
Curve typek-means on 8-week sales ratios (Euclidean vs DTW) · XGBoost multiclass (cross-entropy)Which demand shape a new title will follow
TrajectoryBass diffusion · nonlinear least squaresWeek-by-week curve after release, with interpretable coefficients
MarketProphet · Korean holidays · release-event regressor · ensembleTotal daily sales, decomposed into explainable parts
EvaluationRMSE · wMAPE · mBRSE · weighted accuracy · rolling forecast originsErrors weighted toward the titles and days that matter

Step 1: Build one title-level data set

Every model starts from the same table: one row per title, holding what is known before release and what happened after. VOD purchases are aggregated by title and by week since VOD release. Pre-release features come from title metadata and from public sources — theatrical box office and admissions from KOBIS, Korea's public box-office information system, and audience ratings — together with timing features such as the gap between the cinema release and the VOD release, the release month and the genre.

The post-release side is stored twice: as a total over the selling window, which the title-total model predicts, and as a ratio curve — each week's share of the eight-week total — which separates a title's shape from its size. Two titles with very different revenue can share a shape, and that shared shape is what the curve-type model learns.

data/titles.py
import numpy as np
import pandas as pd

def build_titles(purchases, meta, boxoffice, weeks=8):
    """One row per title: pre-release features, total and weekly sales ratios."""
    p = purchases.merge(meta[["title_id", "vod_release"]], on="title_id")
    p["week"] = (p["date"] - p["vod_release"]).dt.days // 7
    weekly = (p[p["week"].between(0, weeks - 1)]
              .pivot_table(index="title_id", columns="week", values="revenue",
                           aggfunc="sum", fill_value=0))
    total = weekly.sum(axis=1)
    ratios = weekly.div(total, axis=0).add_prefix("r")   # shape only, sums to 1

    t = meta.merge(boxoffice, on="title_id", how="left").set_index("title_id")
    t["gap_days"] = (t["vod_release"] - t["theatrical_release"]).dt.days
    t["log_admissions"] = np.log1p(t["admissions"])
    t["log_total"] = np.log1p(total)
    return t.join(ratios, how="inner")

Simplified. Column names are illustrative; public box-office data supplies admissions and theatrical dates.

Why ratios? Clustering raw revenue curves groups titles by size. Dividing each curve by its total groups them by how demand unfolds, which is the question promotion planning asks.

Step 2: Predict a title's total with tuned tree ensembles

The title-total model is a regression from pre-release features to revenue over the selling window. Revenue is heavy-tailed — a few titles earn most of the money — so the target is modelled on a log scale. Random forest, XGBoost and LightGBM were compared on the same features, and XGBoost was tuned by grid search with five-fold cross-validation choosing the setting.

log ŷ = f_XGB( box office, admissions, ratings, genre, release timing, … )
ŷ    = exp(log ŷ) − 1

Business users asked what drives a forecast as often as they asked for the forecast itself, so the fitted model is explained with permutation importance: each feature is shuffled in turn on held-out titles and the loss in accuracy is measured. Unlike split-count importance, it is computed from the model's predictions and is comparable across features of different types.

models/title_total.py
from sklearn.inspection import permutation_importance
from sklearn.model_selection import GridSearchCV, KFold
from xgboost import XGBRegressor

FEATURES = ["log_admissions", "log_box_office", "rating", "gap_days",
            "release_month", "genre_code"]                    # illustrative

def fit_title_total(train, valid):
    grid = {"max_depth": [3, 4, 6], "learning_rate": [0.03, 0.1],
            "n_estimators": [200, 500], "subsample": [0.8, 1.0]}
    search = GridSearchCV(XGBRegressor(objective="reg:squarederror"), grid,
                          cv=KFold(n_splits=5, shuffle=True, random_state=0),
                          scoring="neg_root_mean_squared_error")
    search.fit(train[FEATURES], train["log_total"])
    model = search.best_estimator_

    # what drives the forecast: loss in accuracy when each feature is shuffled
    imp = permutation_importance(model, valid[FEATURES], valid["log_total"],
                                 n_repeats=20, random_state=0)
    drivers = sorted(zip(FEATURES, imp.importances_mean), key=lambda t: -t[1])
    return model, drivers

# revenue for new titles: np.expm1(model.predict(new_titles[FEATURES]))

Simplified. The grid and feature list are illustrative; random forest and LightGBM were compared the same way.

Step 3: Cluster sales curves into a few demand shapes

A total says nothing about timing. To describe how sales unfold, the eight-week ratio curves of past titles were clustered with k-means under two distances. Euclidean distance compares week w with week w. Dynamic time warping (DTW) lets one curve be stretched or shifted in time against the other before comparing, so a title that peaks in week two and one that peaks in week three can still fall in the same family.

DTW(r, s) = min over warping paths π   Σ_(i,j)∈π (r_i − s_j)²

curve type  c(title) = argmin_k dist(r_title, μ_k),   k = 4
k chosen by within-cluster sum of squares, silhouette and Davies–Bouldin together

No single criterion settles the number of clusters, so within-cluster sum of squares (the elbow), silhouette and Davies–Bouldin scores were read together for both distances; they pointed to four curve types. The cluster centroids are the average shape of each type.

models/curve_types.py
from sklearn.metrics import davies_bouldin_score, silhouette_score
from tslearn.clustering import TimeSeriesKMeans
from tslearn.metrics import cdist_dtw

def compare_k(R, ks=range(2, 8), metric="dtw"):
    """R: (n_titles, 8) weekly sales ratios. Scores every k for one distance."""
    X = R[:, :, None]                                  # tslearn shape (n, T, 1)
    D = cdist_dtw(X) if metric == "dtw" else None
    scores = {}
    for k in ks:
        km = TimeSeriesKMeans(n_clusters=k, metric=metric, n_init=5,
                              random_state=0).fit(X)
        sil = (silhouette_score(D, km.labels_, metric="precomputed")
               if D is not None else silhouette_score(R, km.labels_))
        scores[k] = {"model": km, "wss": km.inertia_, "silhouette": sil,
                     "davies_bouldin": davies_bouldin_score(R, km.labels_)}
    return scores

results = {m: compare_k(R, metric=m) for m in ("euclidean", "dtw")}
km = results[chosen_metric][4]["model"]   # k = 4 from WSS, silhouette and DB
centroids = km.cluster_centers_[:, :, 0]  # (4, 8): average shape of each type

Simplified. Illustrated with tslearn's k-means; Davies–Bouldin is computed on the raw ratios for both distances.

Step 4: Predict a new title's curve type before it sells

Curve types help planning only if they can be known before sales arrive. The clusters differed in box office, admissions and ratings, so a multiclass XGBoost — softmax output, cross-entropy loss — was trained to predict a title's curve type from the same pre-release features as the total model.

P(c | title) = softmax( f_XGB(x) )_c,      loss = − Σ_titles log P(c_true | title)

week w, before release:   revenue_w ≈ ŷ_total · Σ_c P(c | title) · μ_c(w)

The two pre-release models then combine: the predicted total sets the size, and the predicted curve type sets how that total is spread over the eight weeks. In the sketch below the class probabilities weight the cluster centroids, so a title that sits between two shapes gets a curve between them rather than a confident wrong one.

models/curve_type_classifier.py
import numpy as np
from sklearn.model_selection import StratifiedKFold, cross_val_predict
from xgboost import XGBClassifier

def fit_curve_classifier(titles, labels):
    """Curve type (0..k-1) from pre-release features only."""
    clf = XGBClassifier(objective="multi:softprob", eval_metric="mlogloss",
                        max_depth=3, n_estimators=300, learning_rate=0.05)
    oof = cross_val_predict(clf, titles[FEATURES], labels, method="predict_proba",
                            cv=StratifiedKFold(n_splits=5, shuffle=True,
                                               random_state=0))
    clf.fit(titles[FEATURES], labels)
    return clf, oof                         # out-of-fold probabilities to evaluate

def pre_release_curve(clf, total_model, new_titles, centroids):
    """Probability-weighted shape times predicted total: revenue per week."""
    P = clf.predict_proba(new_titles[FEATURES])                # (n, k)
    total = np.expm1(total_model.predict(new_titles[FEATURES]))
    return (P @ centroids) * total[:, None]                    # (n, 8)

Simplified. Weighting centroids by class probability is one way to combine shape and size; settings are illustrative.

Step 5: Fit a Bass curve as sales arrive

After release, the title's own sales become the best evidence. The Bass diffusion model describes how a new product is adopted by a population with three interpretable parameters: m, the eventual number of adopters; p, the coefficient of innovation, for people who buy on their own initiative and front-load sales; and q, the coefficient of imitation, for people who follow others and build sales through word of mouth. Fitted to the first days or weeks by nonlinear least squares, it gives the shape of the rest of the curve.

F(t)     = (1 − e^(−(p+q)t)) / (1 + (q/p)·e^(−(p+q)t))       cumulative adoption share
sales_t  = m · [ F(t) − F(t−1) ]

p large → sales peak at once and decay        q large → sales build, then peak later
models/bass.py
import numpy as np
from scipy.optimize import curve_fit

def bass_cum(t, p, q):
    """Cumulative share of eventual adopters reached by time t."""
    e = np.exp(-(p + q) * t)
    return (1 - e) / (1 + (q / p) * e)

def bass_period(t, m, p, q):
    """Sales in period t: m times the increase in cumulative adoption."""
    return m * (bass_cum(t, p, q) - bass_cum(t - 1, p, q))

def fit_bass(sales, m_guess):
    """Fit m, p, q to the periods seen so far by nonlinear least squares."""
    t = np.arange(1, len(sales) + 1, dtype=float)
    (m, p, q), cov = curve_fit(bass_period, t, sales,
                               p0=[m_guess, 0.03, 0.3],
                               bounds=([0, 1e-4, 1e-3], [np.inf, 1, 2]))
    return m, p, q

def forecast(m, p, q, horizon):
    return bass_period(np.arange(1, horizon + 1, dtype=float), m, p, q)

Simplified. Bounds and starting values are illustrative; the starting market size comes from the title-total model.

Early in a title's life the fit is weakly identified: a fast-decaying blockbuster and a slow-building sleeper look alike for a week. The pre-release total and curve type are what anchor it until enough sales have been seen. In the live model below, the feature-based total enters the Bass fit as a prior on m — strong after seven days, weaker after fourteen.

Why a parametric curve? A flexible model could fit past curves more closely, but a few days of sales cannot pin down a free-form curve, and p and q mean something to the business: a high imitation coefficient marks a title that will reward a second promotion push.

Step 6: Forecast the market with Prophet and known events

The market model forecasts the total of all titles' daily sales, which planners use to set budgets and targets. Prophet was chosen because it is an additive model whose parts can be read and argued with: a piecewise-linear trend, weekly and yearly seasonality, holiday effects and extra regressors. Korean public holidays enter as holiday effects, and a release-event regressor marks the days around the releases of major titles, whose launches lift the whole market.

y(t) = g(t) + s_week(t) + s_year(t) + h_holiday(t) + β · release(t) + ε(t)

Both kinds of event are known in advance — holidays are on the calendar and the release slate is planned — so they can be supplied for the forecast horizon as well as for the history. Adding them was what made the market forecast beat Prophet's default components; the final market forecast was produced as an ensemble.

models/market.py
from prophet import Prophet

def fit_market(daily, releases, horizon_days=28):
    """daily: DataFrame(ds, y) of total sales. releases: DataFrame(ds, release),
    the weight of major-title releases on each day, known in advance."""
    hist = daily.merge(releases, on="ds", how="left").fillna({"release": 0})
    m = Prophet(weekly_seasonality=True, yearly_seasonality=True)
    m.add_country_holidays(country_name="KR")
    m.add_regressor("release")
    m.fit(hist)

    future = m.make_future_dataframe(periods=horizon_days)
    future = future.merge(releases, on="ds", how="left").fillna({"release": 0})
    fc = m.predict(future)
    parts = ["trend", "weekly", "yearly", "holidays", "release"]
    return m, fc[["ds", "yhat", "yhat_lower", "yhat_upper"] + parts]

Simplified. The release weight per day is illustrative; holiday handling uses Prophet's built-in Korean calendar.

Why components? Planners need to argue with a forecast: how much of next month is trend, how much is the holiday, how much is the release slate. A model that returns a number without those parts gets ignored the first time it is wrong.

Step 7: Evaluate with weighted errors over rolling origins

A forecast that is accurate on average but wrong on the biggest titles is wrong where it costs money. Errors were therefore weighted by volume. Weighted MAPE divides the total absolute error by total sales, so a miss on a blockbuster counts for more than the same percentage miss on a niche title. RMSE, mBRSE and a weighted accuracy were reported alongside it.

wMAPE = Σ_i |ŷ_i − y_i| / Σ_i y_i            errors weighted by volume
MAPE  = (1/n) · Σ_i |ŷ_i − y_i| / y_i      every title counts equally

A single train/test split says little about a forecaster that will be refitted every few weeks. The models were instead evaluated over many rolling forecast origins: fit on everything before an origin, forecast the next window, score it, move the origin forward and repeat. For title models the same rule applies to release dates — a model is only ever trained on titles released before the ones it is scored on.

eval/rolling.py
import numpy as np
from prophet.diagnostics import cross_validation

def wmape(y, yhat):
    """Volume-weighted error: total absolute miss over total sales."""
    y, yhat = np.asarray(y), np.asarray(yhat)
    return np.abs(yhat - y).sum() / np.abs(y).sum()

def rolling_market_eval(model, initial="365 days", period="14 days",
                        horizon="28 days"):
    """Refit at many forecast origins; one wMAPE per origin."""
    cv = cross_validation(model, initial=initial, period=period, horizon=horizon)
    per_origin = cv.groupby("cutoff").apply(lambda g: wmape(g["y"], g["yhat"]))
    return per_origin.describe()                 # mean, spread and worst case

def rolling_title_eval(titles, fit, origins):
    """Same idea for title models: train only on titles released before t0."""
    errs = []
    for t0, t1 in zip(origins[:-1], origins[1:]):
        train = titles[titles["vod_release"] < t0]
        test = titles[(titles["vod_release"] >= t0) & (titles["vod_release"] < t1)]
        model = fit(train)
        errs.append(wmape(np.expm1(test["log_total"]),
                          np.expm1(model.predict(test[FEATURES]))))
    return np.array(errs)

Simplified. Window lengths are illustrative; fit is any of the title-level training functions above.

Try the live model

The live model below re-implements simplified versions of the title and market forecasts on a generated catalogue, so you can compare them with the rules of thumb they replace.

Live model, computed in your browser on a generated catalogue of sixty past titles. Left: one new release after its first week — a forecast that scales the catalogue's average curve, against one that combines a feature-based total (a log-linear regression standing in for the tree ensembles) with a Bass curve fitted by grid search to the days seen, using that total as a prior. Right: the market's daily total forecast four weeks ahead by an additive trend, weekly, holiday and release-calendar model fitted by least squares (a simplified stand-in for Prophet), against a seasonal-naive forecast. Curve-type clustering is not reproduced. Switch titles, give the model fourteen days instead of seven, or draw a new catalogue. Open the live model on its own page ↗

Results

Threeforecasts — title total, post-release trajectory, market — from one data set
Acquisitionand free-movie selection informed by expected revenue
Promotiontimed to each title's curve, budgets set against the market forecast

The project provided a data-driven basis for content acquisition, free-movie selection, promotion planning and marketing resource allocation by forecasting both individual content performance and overall market movement.

Splitting the problem is what made the forecasts usable: each model maps to a decision. The title total and its drivers inform what to acquire and what to offer free; the curve type and the Bass coefficients inform when to promote; and the market decomposition separates trend from holidays and the release slate when budgets are set.

Lessons learned

Conclusion

Forecasting for an IPTV movie business is three problems: how much a title will earn, how its sales will unfold, and where the market is going. Treating them separately — tree ensembles for totals, curve clustering and a diffusion model for trajectories, a component model with known events for the market — and evaluating each with weighted errors over rolling origins produced forecasts that matched the decisions they served.

The same structure — a pre-release model for size, a shape model updated as data arrives, and an event-aware model for the aggregate — applies to other launch-driven catalogues such as series, games or new products.

Limitations

About the demo and confidentiality

Titles, admissions, sales figures, the holiday calendar and all errors in the embedded model are generated. No content, sales or customer data from any service, and no internal feature list, hyperparameter setting or evaluation result, appears in this post; code is simplified and written for illustration.

Taehee Lee · Data Scientist / Technical Lead, LG Uplus (2020 – 2023)Problem framing, feature design, modelling of all three forecasts. Demo re-implemented on generated data for this site.