Lailara Promo Incrementality

Accuracy · estimate vs truth

← Back to the Scorecard

How wrong are these numbers?

Every incrementality tool asserts accuracy. This one measures it. The two baselines are scored against known ground truth — the error is shown, by regime, including where it is large. That is the whole claim, and it is a narrow one: this is the error a standard method makes under a realistic, fully-known world. It is not a prediction of the error on your data.

The estimators are provably blind — enforced in code, not promised. An AST gate runs over every estimation file on every push; the generator's own coefficients are banned from the estimation path; and the git history shows both methods frozen and tagged before this page's code first read truth. The blindness claim is scoped exactly there — to the code — and nowhere wider.

Median error, full population

Method 0 · pre-period

26.5%

median absolute error on incremental units

Bias +11.64% · 120 events scored

Method 1 · comparable-store

26.26%

median absolute error on incremental units

Bias +22.02% · 117 events scored

Both methods over-credit promotions — the sign of the bias is positive for each. The comparable-store method, the more defensible one, is more biased, not less: the better baseline does not flatter the promo book, it indicts it further. A demonstration engineered to make the naive method lose would not show this. Large error is the finding, not a blemish to sand off.

The four seeded stories, scored separately

These four events are planted outliers the tool is supposed to surface. They are reported here, apart from the headline median above — the honest denominator is the full population, not the outliers.

StoryMethod 0 errorMethod 1 error
Pure subsidy PRE-0002-35.71%+8.63%
Hero cannibal PRE-0043+4.62%+8.76%
Pantry trap PRE-0056+0.99%+2.19%
Clean winner PRE-0087-15.34%+25.27%

Error by regime

Median absolute error, cut by observed features only — promotion type, depth, season and the like. No cut uses a truth-derived label, and every bucket holds at least five events, so no number reads back to an individual promotion.

Retailer

RetailerMethod 0Method 1
RET-COSTCO13.66%30.74%
RET-KROGER27.15%25.88%
RET-REGIONAL39.04%39.42%
RET-SPROUTS36.58%30.24%
RET-WALMART16.07%20.46%
RET-WHOLEFOODS28.65%15.33%

Promotion type

Promotion typeMethod 0Method 1
BOGO12.92%7.97%
TPR_deep17.51%20.56%
TPR_shallow49.54%38.29%
ad_circular16.53%21.86%
digital_coupon132.18%69.53%
endcap16.09%22.82%

Product line

Product lineMethod 0Method 1
AS33.83%25.61%
DG26.05%23.84%
PS17.78%26.05%
SB16.14%25.27%
SC30.35%29.44%

Discount depth

Discount depthMethod 0Method 1
deep (20–30%)19.95%23.98%
moderate (10–20%)33.02%26.23%
shallow (<10%)27.01%36.95%

Duration

DurationMethod 0Method 1
1 week19.81%25.85%
2 weeks33.02%29.34%
3+ weeks26.86%26.49%

Season

SeasonMethod 0Method 1
Fall35.71%26.49%
Spring19.49%21.21%
Summer15.99%31.25%
Winter49.13%24.21%

Match relaxation — Method 1 only

Method 1 matches comparable stores by region, format class and volume; where the in-format pool is too thin it relaxes to region and volume alone. This cut asks whether the relaxation costs accuracy — and it does.

Match stratumMedian errorBiasEvents
fully relaxed26.26%+22.6%63
mixed26.49%+15.67%53

Synthetic data is the only honest testbed for this. It is at once the only world where truth is knowable and the only world that can be published: real client promotion data can never be shown, by any vendor, at any client count. Anyone claiming to demonstrate accuracy on real client data is either breaching confidentiality or making it up. The methodology and the deliverable are real; the data is synthetic, and that is the point.