Step 02 · Prediction Engine
In progressLarsen & Holt is the same composite as in the Analytics Engine cases. The scenario is illustrative; the industry figures are sourced. The Prediction Engine is in active development; this case describes the workflow it is being built to run.
The deli conversion worked. In the sixty-six stores where the counter came out, the chilled ready-meal range was trading ahead of the old deli’s sales per square metre within a quarter, and the conversion dashboard the trading director pinned on day one had caught the small-format planogram problem in week six. But the new range brought a problem the deli never had. A deli slices to order. A chilled wrap ordered on Tuesday arrives on Thursday and has two days to sell.
Store managers ordered chilled the way convenience retail has always ordered short-life lines: the ERP’s suggested quantity, which is a four-week moving average, adjusted by instinct. Ten weeks in, the numbers showed what that produces. Waste in the converted cabinets ran at eleven percent of units. Availability after five in the afternoon - the commuter hour, when these stores do a third of their chilled sales - ran at eighty-two percent. High waste and empty shelves at the same time is the signature of an ordering signal that is not wrong in one direction but in every direction: too much on a cool Monday, too little on a hot Friday, far too little in a heatwave.
That pattern scales. The UN Environment Programme’s 2024 index attributes 12 percent of the world’s food waste to the retail sector, and the root of much of it is a forecast. How much the method matters at store-and-product level is no longer a matter of opinion: in the M5 competition, run on Walmart’s own unit sales, the top-performing methods were all machine-learning approaches, and all clearly beat the statistical benchmarks. The gap between a moving average and a well-built model is not academic. It is measured in cabinets.
The data team’s estimate for a fresh-demand forecast was twelve weeks to build, plus ongoing ownership no one had budgeted. The replenishment suites that do this for the national grocers are priced for the national grocers. So the wraps kept being ordered by moving average.
The prompt the trading director wrote
Forecast daily unit sales for every chilled ready-meal product in every converted store for the next 14 days. Account for weather, day of week, school holidays, local events and promotions. Tell me how it compares with what we order today.
The Prediction Engine’s first move was to describe what it had been given, and the description shaped everything that followed. Sixty-six stores times forty-eight products is about 3,200 daily series; most were short, because the range was seven months old; many were sparse. Series that short cannot support a model each, so the specification called for one pooled model that lets similar products and similar stores share information - by store format, product family and price point - with calendar, weather forecast, promotions and recent sales as inputs.
It also flagged a trap that hand-built retail forecasts fall into constantly. On days when a cabinet sold out before closing, recorded sales understate demand. The engine used the stock and zero-sales data the Analytics Engine already watched to mark those days, and treated their sales as a floor rather than as the truth. And it tested the way ordering actually works: a rolling eight-week backtest in which every day was forecast as of the order cut-off two days earlier, not with hindsight.
On the held-out weeks, the pooled model’s weighted error was 24 percent against 41 for the moving average, with almost no systematic bias; the moving average consistently over-ordered on Mondays and under-ordered on warm Fridays. Products with less than a month of history were reported separately, at a much weaker 38 percent, rather than being allowed to hide inside the average. When a heatwave arrived in week six, the forecast knew three days ahead, because the weather feed did: salads and wraps up by half, hot pots down.
Units sold per day
Composite scenario; series illustrative. Rolling-origin backtest: each day forecast as of the order cut-off two days earlier.
The other thing the shelves were hiding
The phantom-inventory rule the store operations director wrote in the Analytics Engine - flag a top-500 product with zero sales for a full day while the ERP shows stock - catches fast sellers going dark. It cannot see slow sellers, for which a day without a sale is normal. And the scale of what it cannot see is large: a study of nearly 370,000 records across 37 stores of one retailer found 65 percent of inventory records were inaccurate.
The store operations director asked the next question: predict the probability that each product’s inventory record in each store is wrong by more than three units, trained on last year’s stock-count results. Stock counts are the only ground truth available, so they became the target. The inputs were days since last count, sales speed against the recorded stock, shrink-prone categories, delivery exceptions and store format. The engine trained on sixty stores and tested on the other twenty-five, so the result would reflect stores the model had never seen.
Staff count forty lines a day per store regardless. On the held-out stores, the model’s top forty lines turned out to be wrong 68 percent of the time, against 22 percent for the rotating count schedule and 41 for the zero-sales rule. Same hours on the shop floor; three times the corrections per hour.
Same forty lines a day per store. Three times the errors found per audit hour.
Composite scenario; figures illustrative. Model trained on stock-count results from 60 stores and tested on the other 25.
What it changed, and who owns it
The forecast now drives the suggested chilled order. Store managers can still override it, and every override is logged and scored against what happened. That turned out to be informative in both directions: managers who ordered up ahead of local events were usually right, and the events calendar they were reacting to became a model input. After ten weeks, chilled waste was down from eleven to seven percent and evening availability up from eighty-two to ninety-one.
The two analysts from the analytics story took the new work naturally. The promotions analyst now asks the engine for forecast uplift before a promotion is committed rather than measuring it afterwards. The analyst who owns the alert portfolio added a model-health rule: alert when forecast error for any store cluster moves outside its eight-week range. That rule fired once, when a new sandwich supplier’s deliveries started arriving a day late.
The wraps are still ordered two days early; the difference is that the guess now carries a measured error. But a forecast is not an order. It does not know that tomorrow’s truck is full, that the kitchen can only make so many chicken wraps, or that the back-room chiller holds four cages. Deciding how many to order under those limits is a trade-off with a best answer - and that is a different engine’s job.
Sources
- UNEP, Food Waste Index Report 2024: 1.05B tonnes wasted in 2022; retail 12%.
- Makridakis, Spiliotis & Assimakopoulos, International Journal of Forecasting (2022): M5 Accuracy competition on Walmart sales; top methods pure ML, all beating statistical benchmarks.
- DeHoratius & Raman, Management Science (2008): 65% of ~370,000 inventory records inaccurate.