Supply Chain
Measuring Forecast Accuracy Without Fooling Yourself
Aggregate accuracy always looks better than item-level accuracy, and errors that cancel out report as precision. Pick the measure that can't flatter you.
Forecast accuracy is one of the easiest metrics to accidentally game, because almost every convenient way to summarise it makes the forecast look better than it is. The aggregation level, the choice of denominator and the treatment of offsetting errors each move the number in the flattering direction.
Where the flattery comes from
• Aggregation — total-company accuracy is always better than item-location accuracy, because overs and unders cancel. Only the granular figure is actionable.
• Netting — measuring the signed error lets a +200 and a −200 report as perfect.
• Denominator choice — dividing by the forecast rather than the actual makes over-forecasting look better than it was.
• Time bucket — monthly accuracy hides a forecast that was right in total and wrong every week.
Measure at the level you make decisions. If you replenish by item and location, that is where accuracy has to be measured.
Report bias and accuracy separately
These are different failures with different owners. Accuracy is how far off you were; bias is whether you were consistently off in one direction. A biased forecast is a process problem — sales sandbagging, promotional uplift double-counted, a planner hedging — and it is fixable. An unbiased but inaccurate forecast is a volatility problem, and the answer is inventory policy rather than better forecasting.
Avoid MAPE where demand is lumpy
MAPE divides by the actual, so a single zero-demand period makes it undefined, and small actuals make it explode. For spare parts and slow movers use WAPE or scale the absolute error by average demand instead. Continuing to report MAPE on intermittent items produces numbers nobody can interpret, which is worse than no metric.
Benchmark against doing nothing
Always compute the accuracy of a naive forecast — last period's actual, or the same period last year — alongside your real one. If the sophisticated model does not clearly beat the naive one, that is the finding, and it is more useful than a spuriously precise accuracy percentage.
Connect the metric to a decision
Accuracy is not an end in itself; it is an input to how much safety stock you carry. Reporting the two side by side keeps the conversation honest: an item whose forecast error fell by a third should be carrying less cover, and if it isn't, the improvement has not been banked.
Bias versus accuracy as separate numbers was the unlock for us. We were accurate on average and consistently 8% high on every promoted line, which is two entirely different conversations.
MAPE on intermittent demand is genuinely useless — divide by an actual of zero and the whole month is undefined or infinite. We moved to MAE scaled by average demand and it started behaving.
The naive-forecast benchmark is humbling and everyone should run it once. Our statistical model beat last-period-actual by four points. Four. After two years of investment.