Most demand forecasts are still built in spreadsheets: last year’s volumes, adjusted by experienced planners for what they know about customers, promotions and supply. When someone proposes a machine learning model instead, the discussion quickly turns into a contest between the data science team and the planners. Both sides usually have a point, because the honest answer depends on the item and on how far ahead you are forecasting.
The useful question is not whether the model is better than the spreadsheet. It is where the model is better, by how much, and what planners should do with the time it gives back. You can answer that question with your own history before anything changes in the planning process.
Where the model wins
Machine learning forecasts earn their place under three conditions. The first is enough history: at least two full seasonal cycles at the level you plan, so the model can separate seasonality from trend and from noise. The second is scale: hundreds or thousands of items and locations, more than any team can review one by one each month. The third is stable, measurable drivers, such as price, promotions with a known calendar, seasonality and trading days, recorded consistently in the data.
Under those conditions the model has advantages no planner can match. It looks at every item in every cycle with the same discipline, it notices small shifts in trend early, and it is not anchored on last year’s plan or on the number promised to sales. It is also consistent: the same inputs always produce the same forecast, so every error can be traced back to its cause.
Where judgment wins
A model only knows what is in the data. It has nothing to learn from when a product is new, when an event will not repeat, or when history is too thin or too broken to trust. Launches, a customer lost or won, a competitor’s stock-out, a one-time tender, a plant shutdown and a change of product codes after a system migration all fall into this category.
Planners also hold information that no dataset contains: the call in which a customer said they were changing suppliers, or the promotion a retailer confirmed yesterday. For these items and events a planner’s judgment beats any model, and a good design gives planners the time and the tools to apply it.
How to run a fair test
Most comparisons between a model and the current forecast are unfair to one side. The model is scored on data it has already seen, or the planners are judged on a final, restated forecast rather than the one that drove decisions at the time. A fair test, usually called a back-test, follows five rules:
- Hold out a period the model has never seen. Train on history up to a cut-off date, then forecast the months after it as if they were live.
- Compare at the lag that drives decisions. If production is planned three months out, compare the forecasts made three months before each month, not the latest revision.
- Measure bias as well as error. A forecast can look accurate on average and still run high every month, which is how excess inventory builds up.
- Weight by volume or value. A large miss on a slow item matters less than a small miss on a top seller.
- Report by segment and by horizon. A single average hides the answer you need: where each approach wins.
Add a simple benchmark as well, such as the same month last year with no adjustment. Items where neither the model nor the planners beat it need a simpler process, not a better forecast.
Combine the model and the planner
The organizations that get the most from machine learning do not choose between the model and the planner. The model produces a baseline for every item in every cycle. Planners review the exceptions: items where the model’s range is wide, items whose forecast moved sharply, and the launches and events the model cannot know about. Close in, planners often know things the model cannot, such as an order already in hand; further out, the model’s discipline usually wins. Every change they make is recorded with a reason.
Recording overrides makes it possible to measure what each step adds, a practice known as forecast value added. After a few cycles the pattern is usually clear: adjustments for launches and confirmed promotions improve the forecast, while small changes to steady items rarely do. Planners then spend their time where it pays, and the model carries the rest. Accuracy by horizon, the measure that keeps a rolling forecast honest, does the same job here.
What has to be in place
The model is the smaller part of the work. Sales history has to be clean and defined once: orders or shipments, gross or net of returns, at the level the business plans. Promotions and events need a calendar in the data rather than in someone’s head. The product hierarchy needs launch and discontinuation dates and a link from old items to their successors; otherwise the model treats every replacement as a new product with no history.
The planning cycle changes too. The model refreshes before the monthly demand review, the review starts from the exceptions rather than from the full list of items, and accuracy by segment and horizon is published every month next to the forecast. That is how machine learning becomes part of planning rather than a side project.
Start with a back-test
Before deciding, run the test on your own history: a full year held out, the forecasts that were made at the time, and the model’s forecasts for the same months, compared by segment and by horizon. If the model wins on most of the volume, let it carry that volume and give planners the launches and the events. If it does not, the test usually shows why, and the reason is more often the data than the algorithm.
Share the results with the planners first. They will explain the misses faster than anyone else, and they will trust a model they helped to test far more than one that was handed to them.