Every deck about AI in commercial analytics is built on the same four-rung ladder: descriptive, diagnostic, predictive, prescriptive. It has been on the first slide of that deck for fifteen years. It is not wrong, exactly. It is just too coarse to make a single decision with, and its coarseness is expensive in a specific way: it puts a correlation and a randomised holdout in the same box and calls both of them prescriptive.

This is part one of three. It is the boring part - vocabulary and a frame - and everything practical in parts two and three depends on it holding.

What the work actually is

The useful reframe is not “which model do we build”. FMCG commercial analytics operates over one very high-dimensional cube:

SKU × store/customer × region × week × channel × promo × price × distribution × availability × shopper segment

No human can enumerate that. There are more slices than there are working hours in a year, and the interesting ones are not where you would look. What machine learning is genuinely good at here is not proving causality and it is not chatting with your data. It is the systematic management of commercial attention: finding where performance changed, what it moved with, where an opportunity is sitting unclaimed, roughly how big it is, and which actions are worth testing first.

The largest near-term value of AI/ML in this domain is not “the machine proved causality.” It is that the machine decides what a scarce analyst looks at on Monday morning.

That single move converts the conversation from which model to which decision questions, which is the only version of the conversation a commercial director can participate in.

The Evidence & Action Ladder

Here is the replacement for the four-rung ladder. Ten rungs, each with its own question, its own output, and - the part that matters - its own forbidden overclaim.

#RungThe questionFMCG exampleAllowed phrasingForbidden overclaim
1ObserveWhat happened?Sales −8%, share −1.2pp, availability 82%Observed factAny “why”
2LocalizeWhere exactly?42% of the decline in Retailer A / South / SKU CLocalized issue”Retailer A caused the decline”
3AssociateWhat moved with it?Sales gap associated with availability gap and lost distributionAssociated driver hypothesis”Availability cost us X units”
4SizeWhat order of magnitude?Missed sales ≈ £300kEstimate, conditional on baseline”We will recover £300k”
5PredictWhat is likely next?Out-of-stock risk in two weeksForecasted risk”It will go out of stock”
6SimulateWhat does the model say if X changes?Discount 15% → margin £1.3mModel-implied scenario”Raising availability will yield Y”
7RecommendWhat do we do first?Fix the top 100 availability issuesRecommended action under a rule”This action returns exactly £X”
8ValidateDid the action work?Post-intervention recovery vs baselineValidated action-
9Causal estimateWhat was the incremental effect?Difference-in-differences promo uplift +6%Estimated under stated assumptions”Proven, universally”
10OptimizeHow do we allocate a scarce resource?Field team visits 80 stores - which?Optimized portfolio”The optimizer created the effect”

Two things about this table are worth more than the table itself.

Localization and association are not warm-up exercises. In most published maturity models, rungs 2 and 3 are the tedious prelude before the real work begins. In FMCG they are where the majority of realisable value lives, because the cube is large and nobody is currently searching it systematically. A team with no data science function should be aiming to run rungs 1–7 well, permanently, rather than aiming to reach rung 9 badly.

The forbidden-overclaim column is the actual product. A finding without an evidence label is not a finding, it is a rumour with a chart attached. The label is what stops a rung-3 association being presented to a board as a rung-9 effect, which is the single most common way analytics loses its credibility - not by being wrong, but by being right at a lower rung than it claimed.

One open question, two directions of travel

Above all ten rungs sits one question that is never answered by a dashboard:

What is the state of the business?

It decomposes into tracks, and the tracks split on a distinction that gets collapsed constantly:

TrackHow you enterMain questionOutput
State of businessOpen scanWhat is going on at all?Executive radar: risks and openings
Issue investigationA known event, top-downWhy did sales / share / margin fall?Issue pack + next action
Opportunity finderOpen search, bottom-upWhere is the upside when nothing is broken?Opportunity backlog, whitespace list
Missed sales finderDemand existed, was not servedWhere did we underserve on availability or listing?Missed-sales action queue
Promo / RGM decision supportA decision is dueWhich promo, price or pack?Scenario + recommendation
Causal validationEscalation for expensive decisionsCan we say X caused Y?Causal estimate memo

Issue investigation is top-down: the event is known, and you are narrowing. Opportunity finding is bottom-up: there is no event, and you are scanning a universe. These are different modes of thinking with different failure modes, and merging them into one process - which most “AI insights” tooling does - produces something that is mediocre at both. Part two takes each apart.

The words people keep swapping

Half the arguments in this field are vocabulary arguments that nobody has noticed are vocabulary arguments.

TermWhat it meansWhat it does not mean
LocalizationFinding where a change is concentratedIt does not identify a driver
Slice findingAutomated search across dimension combinations for unusual behaviourA slice is a subgroup, not a cause
AssociationA variable is statistically linked to the target, or helps explain or predict itNot a causal effect
Driver diagnosisBusiness interpretation of associated signals into a likely explanationStill a hypothesis
PredictionEstimating a future value or risk under current patternsIt does not choose an action
Scenario / what-ifRecomputing a model with a changed inputThe coefficient still has to come from somewhere
PrescriptionA recommended actionCan be strong or very weak on evidence
Causal effectThe estimated effect of an intervention, do(X)Requires a design, not just a model
OptimizationChoosing the best set of actions under constraintsIt does not create the coefficients it optimises over

If a team adopts nothing else from this piece, adopt this table and put it on a wall. Most of the expensive mistakes downstream are one of these nine confusions wearing a suit.

”Prescriptive” is a spectrum, not a rung

This is the correction I would most like to see propagate. In vendor language, prescriptive analytics absorbs forecasting, what-if, regression, scenarios, recommendations, optimization and causal prescription into one word. In practice it is a spectrum with an order of magnitude of rigour between its ends:

Kind of prescriptionRests onExampleRigour
Rule-basedA business ruleAvailability <85% on a top SKU → fix itLow, but genuinely practical
Threshold / exceptionAn SLA or limitDistribution −5pp → check the listingLow
Association-basedCorrelation, regression, SHAPSales gap associated with availability gap → check availabilityMedium; good enough for triage
Forecast-basedPredicted riskOut-of-stock risk → replenishMedium
Scenario / what-ifA coefficient plus a formulaDiscount 15% → best marginEntirely dependent on the coefficient
Simulation-basedResponse curves plus assumptionsSimulating a promo calendarMedium to high
Score-based prioritisationValue × confidence × feasibilityTop 100 store-SKU actionsVery practical
Optimization-basedExpected value under constraintsAllocating field-force visitsHigh, for decision support
Causal-effect-basedDiD, BSTS, DML, holdoutAvailability fix → +6%High
Experiment-backedRandomised controlled testRollout after a holdoutStrongest available
AdaptiveA feedback loopThe system learns from each actionMature

The frame that makes this usable:

Prescription = action logic + expected value + constraints + evidence label

Drop the fourth term and the whole thing degrades into advice. The failure starts exactly where a recommendation derived from an association gets sold as a proven incremental effect - and it is usually sold that way not out of dishonesty but because the vocabulary offered no way to say the difference out loud.

Where the coefficient comes from

This is the layer I missed on my first pass through the problem, and it turned out to be the load-bearing one. Any what-if or any prescription with a number attached requires a response coefficient: how much does Y move when X moves. And the crucial realisation is that the method does not determine the strength of the claim - the design of the question and the origin of the coefficient do.

Coefficient sourceExampleWhat to call the outputRisk
Expert assumption”+10% discount → +20% units”Assumption-based scenarioSubjective
Historical averageMean uplift of similar past promosHistorical scenarioConfounded by season, support, stock
CorrelationAvailability and sales move togetherAssociation-based hypothesisConfounding, reverse causality
Regression coefficientSales ~ availability + price + promo + controlsModel-based associationNot causal without a design; multicollinearity
Predictive ML / SHAPModel predicts, SHAP highlights featuresPredictive scenarioFeature importance ≠ causal effect
Quasi-experimentTreated vs comparable control, before and afterCausal estimate under assumptionsNeeds a real design
Randomised holdoutRandom test and control storesExperimental effectExpensive, slow, external validity
Vendor benchmarkExternal elasticityBenchmark scenarioMay not transfer
Feedback loopReal outcomes of past actionsAdaptive estimateRequires clean action logging

Read down that table and the point lands: a what-if is not evidence. Its quality is entirely inherited from the row it drew its coefficient from. The same regression can be an association, a prediction, a scenario or a causal estimate depending on how the question was designed - and nothing in the output distinguishes the four. Only the design does, and only a human records it.

What part one buys you

Nothing yet. That is the honest answer. This part builds no model, finds no opportunity and saves no money.

What it does is make the next two parts possible to write without lying. Part two builds the engine - the three finders, the grain they need, and the exact point where a BI tool stops being able to help. Part three asks who has to be in the room for any of it, what it costs, and why the first project should not be a vendor implementation.



AI in FMCG analytics, part two — the engine