The structural reason your dashboard missed the regime change.

A forecast is a claim about a process. Most forecasts in this industry are claims about the wrong process, and they’re confident about it, which is the expensive part.

There are two ways to end up there. They look nothing alike and they fail identically.

The complaint turns up everywhere, always from people who sell data or prediction - so take the sales interest out and see what’s left. Charles Chase at SAS, on consumption-based forecasting, says the problem is forecasting shipments and calling it demand. Miro Dimitrov at NWO.ai, on demand drivers, says the problem is forecasting on first-party data and calling it the market. WorldQuant Predictive went further and built a platform that refuses to integrate with client systems at all. Kim Cox at NielsenIQ, on omnichannel measurement, says retail point-of-sale has itself stopped being a full view, because purchase moved into channels that don’t report.

Nobody connects them. They’re the same failure.

One: you’re forecasting supply

Forecast shipments into the channel and you’re measuring what you sent, not what people consumed. In a stable period the two track each other closely enough that nobody notices. In a shock they come apart immediately, and the model has no way of knowing which of the two it was ever describing.

The last five years made this visible twice - the pandemic, then the inflation shocks. What broke wasn’t the model. It was the assumption underneath it, that a shipment is a proxy for a purchase, and a model can’t report the failure of an assumption nobody told it it was making.

The fix is structural rather than statistical. Put a consumption signal in front of the shipment model and forecast the thing you actually care about further up the chain.

Two: you’re forecasting your own customers

Forecast on first-party data only and you’re measuring people who already found you. They’re your customers. Everything the model can see is downstream of that. It’s an echo chamber with excellent data hygiene.

This one is harder to spot, precisely because first-party data is usually the cleanest data a company has. Quality isn’t the problem. Selection is. Nothing in a completeness check, a null-rate report or a schema test detects a population that was chosen for you before you started.

Refusing to integrate with internal systems is the same argument as a product decision: internal data isn’t an asset to add, it’s a bias to keep out. And the point about point-of-sale is the same failure one level up - a source that used to be the market has quietly become a sample, and nothing in the data announced the change.

The rule underneath

Two failures that look nothing alike, one cause. What a model can observe decides what it’s capable of noticing.

Which gives the rule I actually work by, and which none of the four states:

Every production model needs a named signal by which it learns the regime has changed - and that signal can’t come from the same source as its training data.

A model trained on shipments can’t be monitored on shipments. A model trained on first-party behaviour can’t be monitored on first-party behaviour. In both cases the monitoring inherits the blindness of the training, and the model reports that everything is fine right up to the point where it isn’t.

ModelTrained onIndependent regime signal
Demand forecastShipments to channelSyndicated consumption, POS
Category growthOwn POSPanel, receipts, search
ChurnFirst-party behaviourCategory penetration, external switching
Innovation demandOwn launchesScientific literature, venture funding

The fourth row is the one people push back on, and it comes from NWO’s horizon rule: a question about the next twelve months and a question about the next decade need different sources. Most teams run both through the same corpus because it’s the corpus they have. It’s a cheap thing to fix and it stops a five-year question being answered with last quarter’s data.

From my own work

I’ve run this failure from the inside.

At a wine-and-spirits company the demand plan forecast shipments into distributors and key accounts and called the result demand. Granular, confident - and built on sell-in. What people actually consumed wasn’t in it, because sell-out data availability was . The model was trained and monitored on the one number it could see.

Nothing looked wrong until the trade was full of extra stock and some brands had stopped moving, and the forecast had no way to say so. Every number it watched was a number it had already sold. It reported a healthy regime right up to the warehouses saying otherwise.

The fix wasn’t a better model. It was a different thing to forecast. We rebuilt the plan around sell-out - market growth, promotions, competition, seasonality - and turned sell-in into arithmetic on top: stock level, stock policy, done. The signal the model needed to notice the regime had changed was the one that had never been in the room, and we had to go and get it before it could be used. That’s the rule from above, and it cost more in conversations than in code.

Accuracy went up and forecasts landed much earlier. A forecast that can’t see consumption can’t be fixed by any amount of statistics aimed at shipments.

The cheap version

You don’t need a programme for this.

Take your production models. For each one, write down the single signal that would tell you it has stopped describing reality - and check that the signal comes from somewhere else.

If you can’t name one, you don’t have a monitored model. You have a model and a habit.