Turning daily sales forecasts into a month-ahead traffic-light signal for target risk
A walkthrough of how daily sales forecasts become a signal a business team can act on: separating what can be forecast from what should only be monitored, modelling the calendar, carrying the current level forward, turning a month-end projection into a probability of hitting target and showing it as a traffic light — with simplified code for each step.
Monthly targets are usually checked after the books close, when the only thing left to do is explain the variance. The common early read is a run-rate: month-to-date sales scaled to the full month. It assumes every day is alike, and in a food business they are not — weekends, holiday weeks and the push of orders into the last working days make the early weeks look weak almost every month. A run-rate signal therefore raises alarms in months that end fine, and teams learn to ignore it.
In this post we walk through the predictive-management work for CJ CheilJedang. We scoped the indicators with domain experts, separating predictable areas from highly uncertain ones; integrated sales, cost and expense indicators; built daily sales forecasting models on engineered calendar features; projected each month's total from actuals plus forecasts; converted it into a probability of reaching the target; and showed that as a traffic light on a dashboard. The daily sales model reached 97.6% prediction accuracy for the Korea Food business unit, and teams could see monthly target risk about a month in advance. The specific model, persistence and probability formulas shown below are those of the live model, a re-implementation on a generated business unit.
Solution overview
The system has a data path that integrates business indicators into one daily view, a forecasting path that refreshes daily sales forecasts as new actuals arrive, and a signal path that projects each month's total, compares it with the target and publishes a traffic light. Forecasting covers the predictable areas; highly uncertain areas are monitored rather than forecast.
Predictable vs uncertain
One daily view of the business
Calendar and business features
Daily sales models
Month-end total and P(hit)
Traffic light on a dashboard
The numbered steps in the diagram:
- With domain experts, the data and modelling scope were defined and indicators split into predictable areas to forecast and highly uncertain ones to monitor.
- Sales, cost of sales, selling and administrative expenses and other business indicators were collected and integrated into one view.
- Feature engineering encodes the calendar the business runs on: weekdays, month-end working days, holiday weeks and the days before them, and season.
- Daily sales forecasting models predict each remaining day; the live model uses a calendar regression per product line with a persistent level.
- The month's total is projected as actuals to date plus forecasts for the remaining days, with an uncertainty that shrinks as the month fills, giving a probability of reaching the target.
- The probability is shown as a traffic light on a dashboard for business users, about a month before the books close.
Technology stack
| Layer | Technology | What it does here |
|---|---|---|
| Scope | Scoping with domain experts | Which indicators to forecast and which to monitor |
| Data | Integrated sales, cost of sales, SG&A and other indicators | One daily view of the business unit |
| Features | Calendar and business feature engineering | Weekday rhythm, month-end push, holidays, season |
| Forecasting | Daily sales forecasting models | 97.6% prediction accuracy, Korea Food business unit |
| Risk | Month-end projection with uncertainty | Probability of reaching each monthly target |
| Delivery | Traffic-light signal system · dashboard | Target risk about a month before closing |
| Live model (this page) | Per-line calendar regression · persistence φ · normal approximation | Browser re-implementation on a generated business unit |
Step 1: Decide what to forecast and what to monitor
Not every line of a P&L can be forecast. Working with domain experts, we defined the data and modelling scope, collected and processed sales, cost of sales, selling and administrative expenses and other business indicators, and made the key decision: separate the predictable areas from the highly uncertain ones.
| Predictable areas | Highly uncertain areas | |
|---|---|---|
| Illustrative example | Sales with a stable calendar structure | One-off cost items |
| Treatment | Forecast daily, project the month | Monitored, not forecast |
| Role in the signal | Drives green, amber or red | Kept out of the forecast |
Why draw the line? A signal that mixes forecastable and unforecastable items turns red for reasons nobody can act on. Keeping the forecast to what can be predicted means a red light says the model expects a miss — not that something unpredictable might happen.
Step 2: Encode the calendar the business runs on
Daily sales of a food business follow the calendar more than a trend: weekday ordering, a push of orders into the last working days of the month, holiday weeks such as Seollal and Chuseok with gift buying before them, and seasonality. In the project, the daily models were built on engineered calendar and business features. The live model uses the calendar part, as one table with a row per day:
- Weekday dummies — weekends fall much further for lines sold to business customers than for retail lines.
- Month-end — the last three working days of each month.
- Holiday and pre-holiday — the holiday days, and the ten days before them.
- Trend and season — a linear trend and two annual Fourier harmonics.
import numpy as np
import pandas as pd
def calendar_table(start, end, holidays, pre_days=10, month_end_days=3):
"""One row per day with the calendar features of the live model."""
cal = pd.DataFrame({"date": pd.date_range(start, end, freq="D")})
cal["dow"] = cal["date"].dt.dayofweek
cal["holiday"] = cal["date"].isin(pd.to_datetime(holidays))
# days shortly before a holiday: gift buying and stocking up
hol = cal["holiday"].astype(float)
ahead = hol[::-1].rolling(pre_days, min_periods=1).max()[::-1].shift(-1, fill_value=0)
cal["pre_holiday"] = (ahead > 0) & ~cal["holiday"]
# last working days of each month: the month-end order push
work = cal[(cal["dow"] < 5) & ~cal["holiday"]]
last = work.groupby(work["date"].dt.to_period("M")).tail(month_end_days)
cal["month_end"] = cal["date"].isin(last["date"])
# slow trend and two annual harmonics
cal["trend"] = (cal["date"] - cal["date"].iloc[0]).dt.days / 365
w = 2 * np.pi * cal["date"].dt.dayofyear / 365.25
for k in (1, 2):
cal[f"sin{k}"], cal[f"cos{k}"] = np.sin(k * w), np.cos(k * w)
return calSimplified. Holiday dates come from the public calendar; window lengths are the live model's.
Step 3: Fit a daily calendar regression per product line
In the project, the daily sales forecasting models reached 97.6% prediction accuracy for the Korea Food business unit's sales. The live model uses the simplest model that captures the calendar: a least-squares regression of log daily sales on the calendar table, fitted separately for each product line. Lines differ in shape — a line sold mainly to business customers drops at weekends and peaks at month-end, a gift line peaks before holidays — so effects are not shared. On the log scale effects are multiplicative: a holiday removes a share of sales, not a fixed amount.
log y_(l,t) = x_tᵀ β_l + r_(l,t) x_t = [ 1, trend, weekday (6), month-end, holiday, pre-holiday, annual sin/cos (4) ] 15 terms β_l least squares on the first two years, per product line l r_(l,t) residual: what the calendar does not explain
import numpy as np
import pandas as pd
import statsmodels.api as sm
TERMS = ["trend", "month_end", "holiday", "pre_holiday", "sin1", "cos1", "sin2", "cos2"]
def design(cal: pd.DataFrame) -> pd.DataFrame:
dow = pd.get_dummies(cal["dow"], prefix="dow", drop_first=True, dtype=float)
X = cal[TERMS].astype(float).join(dow)
return sm.add_constant(X).set_index(cal["date"])
def fit_lines(sales: pd.DataFrame, cal: pd.DataFrame, train_end):
"""sales: one column per product line, indexed by the dates of `cal`."""
X = design(cal)
train = X.index < train_end
fits, resid = {}, {}
for line in sales.columns:
y = np.log(sales[line])
fits[line] = sm.OLS(y[train], X[train]).fit()
resid[line] = y - fits[line].predict(X)
return fits, X, pd.DataFrame(resid)Simplified. One ordinary least-squares fit per product line; the residuals feed the next step.
The residual is where the current state of the business shows. If a line has run above its calendar for the last four weeks, that is a level, and levels tend to persist.
Step 4: Carry the current level forward with persistence φ
Calendar effects are known in advance; the level is not. The live model takes each line's current level as its mean residual over the last 28 days, holidays excluded, and lets it fade for months further ahead. How fast is estimated from history: φ is the lag-one autocorrelation of monthly mean residuals, pooled across lines and clipped to the range 0–0.9.
level_l(d) = mean of r_(l,t) over the 28 days up to d, holidays excluded ŷ_(l,t) = exp( x_tᵀ β_l + level_l(d) · φ^k + s²/2 ), k = months between d and t φ = Σ_m ρ_(m−1) ρ_m / Σ_m ρ_(m−1)², ρ_m = monthly mean residual, 0 ≤ φ ≤ 0.9 s = within-month s.d. of daily residuals; s²/2 corrects the log-normal mean
import numpy as np
import pandas as pd
def persistence(resid: pd.DataFrame, cap=0.9) -> float:
"""phi: lag-1 autocorrelation of monthly mean residuals, pooled over lines."""
monthly = resid.groupby(resid.index.to_period("M")).mean().to_numpy()
prev, cur = monthly[:-1], monthly[1:]
return float(np.clip((prev * cur).sum() / (prev ** 2).sum(), 0.0, cap))
def current_level(resid: pd.DataFrame, holiday: pd.Series, asof, window=28):
"""Mean residual per line over the last `window` days, holidays excluded."""
recent = resid.loc[:asof].tail(window)
return recent[~holiday.loc[recent.index]].mean()
def forecast_total(fits, X, level, phi, s, asof, days):
"""Business-unit total per day; each line's level fades as phi ** k."""
k = np.maximum(0, np.round((days - asof).days.to_numpy() / 30.4)) # months ahead
out = {}
for line, fit in fits.items():
base = fit.predict(X.loc[days])
out[line] = np.exp(base + level[line] * phi ** k + s ** 2 / 2)
return pd.DataFrame(out, index=days).sum(axis=1)Simplified. fits and X come from Step 3; the business-unit total is the sum over product lines.
Why estimate φ? Carrying the level unchanged overstates how much this month says about the next; dropping it throws that information away. One estimated φ lets the data decide, and it is what lets the model speak about a month that has not started yet.
Step 5: Project the month and convert it into P(hit target)
On any day d, the month's total is what has been booked plus what the model expects for the remaining days. Its uncertainty has two sources: day-to-day noise, which largely averages out over a month, and a month-level swing in the business unit's total, which does not. Both apply only to the share of the month still to come, so the uncertainty shrinks as actuals replace forecasts.
T̂_m(d) = Σ_(t ≤ d) y_t + Σ_(t > d, t ∈ m) ŷ_t w = Σ_(t > d) ŷ_t / T̂_m(d) share of the month still to come σ(d) = w · √( σ_u² · g + s²/28 ) g = 1 + φ² for a month not yet started P(hit) = Φ( log( T̂_m(d) / target_m ) / σ(d) )
σ_u is the standard deviation of month-to-month surprises in the business unit's total, estimated from an AR(1) on monthly log ratios of actual to calendar-fitted sales. In the live model the light is green at P(hit) ≥ 0.65, red below 0.35 and amber in between; the project's thresholds are not given here.
import numpy as np
from scipy.stats import norm
def month_projection(actual_to_date, forecast_rest, target,
sigma_u, s, phi, months_ahead, g_current=0.6):
"""Projected month total and P(total >= target), log-normal approximation."""
total = actual_to_date + forecast_rest
w = forecast_rest / total # share of the month still to come
g = 1 + phi ** 2 if months_ahead >= 1 else g_current
sigma = max(w * np.sqrt(sigma_u ** 2 * g + s ** 2 / 28), 1e-4)
p_hit = norm.cdf(np.log(total / target) / sigma)
return total, sigma, p_hit
def light(p_hit, p_hi=0.65, p_lo=0.35):
return "green" if p_hit >= p_hi else "amber" if p_hit >= p_lo else "red"Simplified. g_current is a hand-set factor for the month in progress in the live model; the thresholds are the live model's.
Step 6: Show a traffic light and test it against the run-rate
The probability is what the model knows; the traffic light is what a business team reads. In the project, the signal system and dashboard were designed so that business users could understand target-achievement risk quickly and respond.
To check that the light is useful, the live model replays the third year month by month against a run-rate signal: green if month-to-date sales scaled to the full month reach the target, amber within 3% of it, red otherwise. It counts the missed months flagged red on the 10th, the months that hit but were flagged red, and how many days before the close a missed month's light turned red and stayed red.
import pandas as pd
def run_rate_light(mtd, days_elapsed, days_in_month, target, amber=0.97):
projection = mtd * days_in_month / days_elapsed
if projection >= target:
return "green"
return "amber" if projection >= amber * target else "red"
def red_lead_days(daily_lights: pd.Series) -> int:
"""Days before close during which the light was red and stayed red."""
n = 0
for colour in daily_lights.iloc[::-1]:
if colour != "red":
break
n += 1
return n
def score_year(months: pd.DataFrame, light_col: str, lead_col: str) -> pd.Series:
"""months: one row per test month with `missed` and one method's columns."""
missed, hit = months[months["missed"]], months[~months["missed"]]
return pd.Series({
"misses_flagged_red": int((missed[light_col] == "red").sum()),
"missed_months": len(missed),
"false_red": int((hit[light_col] == "red").sum()),
"mean_lead_days": missed[lead_col].mean(),
})Simplified. Lights are recorded for every day from a month before the month starts to its last day.
The run-rate's weakness is structural: with a month-end push, the first ten days are always below a straight-line share of the month, so the run-rate turns red in months that end fine. The daily model already expects the push.
| Component | In the project | In the live model |
|---|---|---|
| Models | Daily sales forecasting models (time-series, engineered features) | A calendar regression per product line with a persistent level |
| Signal | Traffic-light signal system and dashboard | The same idea, thresholds at 0.65 and 0.35 |
| Data | Korea Food business-unit sales and other business indicators | Six generated product lines, amounts as % of target |
Try the live model
The live model below replays a generated test year month by month, so you can compare the run-rate signal with the daily model's probability of hitting the target — on the 10th, at the start of the month and a month ahead.
Live model, computed in your browser on a generated business unit of six product lines over three years, with a weekday rhythm, an end-of-month push, Seollal and Chuseok weeks, seasonality and persistent monthly swings; amounts are shown as a percentage of target, and targets are last year's month plus 5%. Left: the month-to-date run-rate and its signal. Right: a daily model fitted on two years, its projection and probability of hitting the target, and its signal from a month before the month starts. The table compares both across the test year. The features of the demo business are invented to show why a calendar-aware forecast beats a run-rate. Open the live model on its own page ↗
Results
The daily sales model achieved 97.6% prediction accuracy for CJ CheilJedang's Korea Food business unit sales. The signal system enabled business teams to identify monthly target risk about one month in advance, shifting management from after-the-fact review to proactive decision making.
Two design choices sit behind the signal: the forecast is limited to predictable areas, so a red light has a forecastable cause, and the output is a probability shown as a light, so business users can read it at a glance.
Lessons learned
- Decide what not to forecast. Separating predictable areas from highly uncertain ones kept the signal interpretable: a red light means the model expects a miss.
- Model the calendar first. Weekday rhythm, month-end orders and holiday weeks shape every month, and they are exactly what fools a run-rate.
- Separate the calendar from the level. Calendar effects are known in advance; the current level fades with time. Modelling them separately lets a forecast speak about next month as well as this one.
- Report a probability, show a colour. Business users act on a light; the probability behind it keeps the light consistent as uncertainty shrinks through the month.
Conclusion
Month-end variance reports explain the past. Scoping the forecastable indicators with domain experts, integrating the business indicators, forecasting daily sales with the calendar built in and converting the month's projection into a probability and a traffic light gave CJ CheilJedang's business teams a view of target risk about a month ahead, with 97.6% daily sales prediction accuracy for the Korea Food business unit.
The pattern — forecast what is predictable, project the period's total, state the risk as a probability and show it simply — carries over to other targets tracked monthly, such as volumes or service levels.
Limitations
- The signal is only as good as the target: targets set without regard to what is achievable produce red lights that are correct and unhelpful.
- Shocks the calendar does not know about — a recall, a sudden price change — show up in the forecast only once they show up in sales.
- Areas that are monitored rather than forecast are outside the light by design; a month can miss for reasons the signal does not cover.
- The live model's business unit, lines and targets are generated; it uses a per-line calendar regression with one persistence parameter and a log-normal approximation for the month total, and its accuracy figures describe that generated business, not CJ CheilJedang's.
About the demo and confidentiality
Product lines, sales, holidays and targets in the embedded model are generated, and amounts are shown only as a percentage of target. No sales, cost, target or organizational data from CJ CheilJedang, and no model configuration or signal thresholds used in the project, appear in this post; code is simplified and written for illustration.