Accuracy · estimate vs truth
← Back to the ScorecardHow wrong are these numbers?
Every incrementality tool asserts accuracy. This one measures it. The two baselines are scored against known ground truth — the error is shown, by regime, including where it is large. That is the whole claim, and it is a narrow one: this is the error a standard method makes under a realistic, fully-known world. It is not a prediction of the error on your data.
The estimators are provably blind — enforced in code, not promised. An AST gate runs over every estimation file on every push; the generator's own coefficients are banned from the estimation path; and the git history shows both methods frozen and tagged before this page's code first read truth. The blindness claim is scoped exactly there — to the code — and nowhere wider.
Median error, full population
Method 0 · pre-period
26.5%
median absolute error on incremental units
Bias +11.64% · 120 events scored
Method 1 · comparable-store
26.26%
median absolute error on incremental units
Bias +22.02% · 117 events scored
Both methods over-credit promotions — the sign of the bias is positive for each. The comparable-store method, the more defensible one, is more biased, not less: the better baseline does not flatter the promo book, it indicts it further. A demonstration engineered to make the naive method lose would not show this. Large error is the finding, not a blemish to sand off.
The four seeded stories, scored separately
These four events are planted outliers the tool is supposed to surface. They are reported here, apart from the headline median above — the honest denominator is the full population, not the outliers.
| Story | Method 0 error | Method 1 error |
|---|---|---|
| Pure subsidy PRE-0002 | -35.71% | +8.63% |
| Hero cannibal PRE-0043 | +4.62% | +8.76% |
| Pantry trap PRE-0056 | +0.99% | +2.19% |
| Clean winner PRE-0087 | -15.34% | +25.27% |
Error by regime
Median absolute error, cut by observed features only — promotion type, depth, season and the like. No cut uses a truth-derived label, and every bucket holds at least five events, so no number reads back to an individual promotion.
Retailer
| Retailer | Method 0 | Method 1 |
|---|---|---|
| RET-COSTCO | 13.66% | 30.74% |
| RET-KROGER | 27.15% | 25.88% |
| RET-REGIONAL | 39.04% | 39.42% |
| RET-SPROUTS | 36.58% | 30.24% |
| RET-WALMART | 16.07% | 20.46% |
| RET-WHOLEFOODS | 28.65% | 15.33% |
Promotion type
| Promotion type | Method 0 | Method 1 |
|---|---|---|
| BOGO | 12.92% | 7.97% |
| TPR_deep | 17.51% | 20.56% |
| TPR_shallow | 49.54% | 38.29% |
| ad_circular | 16.53% | 21.86% |
| digital_coupon | 132.18% | 69.53% |
| endcap | 16.09% | 22.82% |
Product line
| Product line | Method 0 | Method 1 |
|---|---|---|
| AS | 33.83% | 25.61% |
| DG | 26.05% | 23.84% |
| PS | 17.78% | 26.05% |
| SB | 16.14% | 25.27% |
| SC | 30.35% | 29.44% |
Discount depth
| Discount depth | Method 0 | Method 1 |
|---|---|---|
| deep (20–30%) | 19.95% | 23.98% |
| moderate (10–20%) | 33.02% | 26.23% |
| shallow (<10%) | 27.01% | 36.95% |
Duration
| Duration | Method 0 | Method 1 |
|---|---|---|
| 1 week | 19.81% | 25.85% |
| 2 weeks | 33.02% | 29.34% |
| 3+ weeks | 26.86% | 26.49% |
Season
| Season | Method 0 | Method 1 |
|---|---|---|
| Fall | 35.71% | 26.49% |
| Spring | 19.49% | 21.21% |
| Summer | 15.99% | 31.25% |
| Winter | 49.13% | 24.21% |
Match relaxation — Method 1 only
Method 1 matches comparable stores by region, format class and volume; where the in-format pool is too thin it relaxes to region and volume alone. This cut asks whether the relaxation costs accuracy — and it does.
| Match stratum | Median error | Bias | Events |
|---|---|---|---|
| fully relaxed | 26.26% | +22.6% | 63 |
| mixed | 26.49% | +15.67% | 53 |
Synthetic data is the only honest testbed for this. It is at once the only world where truth is knowable and the only world that can be published: real client promotion data can never be shown, by any vendor, at any client count. Anyone claiming to demonstrate accuracy on real client data is either breaching confidentiality or making it up. The methodology and the deliverable are real; the data is synthetic, and that is the point.