Attribution Is Not Incrementality
Attribution assigns credit for conversions you observed. Incrementality estimates the ones that would not have happened otherwise. The gap between them is where budgets go wrong.
Every marketing team has had this meeting. Paid social reports a 4.2x return. The finance team points out that revenue is flat. Someone suggests the tracking is broken. Someone else suggests finance is looking at the wrong number.
Usually nothing is broken. The two sides are measuring different things, and both are measuring them correctly.
The two questions
Attribution answers: of the conversions I observed, which touchpoints came before them, and how should I divide credit among those touchpoints?
Incrementality answers: how many of those conversions would not have happened if this campaign had never run?
These sound close enough to be interchangeable. They are not, and the reason is structural rather than technical. Attribution works entirely on conversions that happened. It has no access to a world where the campaign did not run, because that world is not in the data. Attribution divides a total you observed. Incrementality estimates the difference between two totals, only one of which you ever saw.
This is why the debate about first-touch versus last-touch versus data-driven attribution never resolves. Those are different rules for splitting the same pie. None of them is wrong at splitting it. None of them tells you how big the pie would have been without you.
What the gap looks like
Here is a quarter of a mid-size ecommerce program, with every channel’s last-touch credit next to what a holdout measured for the same channel in the same window.
- Last-touch credited
- Measured incremental
Show the data
| Conversions per quarter | Last-touch credited | Measured incremental | Gap |
|---|---|---|---|
| Branded search | 12k | 1.5k | 11k |
| Retargeting | 8.9k | 1.6k | 7.3k |
| Email lifecycle | 7.3k | 2.4k | 4.9k |
| Non-brand search | 6.2k | 3.3k | 2.8k |
| Paid social prospecting | 4.8k | 3.2k | 1.6k |
| Online video | 1.2k | 2.2k | -940 |
Across the program, last-touch credits 40,860 conversions. The experiments measure 14,256, about 35% of the credited total.
The overall ratio matters less than how differently it lands by channel, and in which direction.
Why the gap is largest near the purchase
Look at the ordering. Branded search and retargeting sit at the top of the gap; online video sits at the bottom, and is the only channel where last-touch under-credits.
That ordering is not specific to this dataset. It falls out of how last-touch works.
Last-touch credits whatever a person interacted with most recently before converting. That favors channels positioned near the end of the journey, and channels near the end of the journey are reaching people who have already shown intent.
Someone who types your brand name into a search engine has decided to find you. The paid ad above the organic result gets the click and the credit. Many of those people would have scrolled two centimeters and clicked the organic link. Nothing is broken here. The click happened and the ad did precede the conversion. The ad just changed the outcome for a much smaller group than the click count suggests.
Retargeting is the sharper version of the same problem, because its audience is defined by prior intent. You select the people most likely to convert, then advertise to them. Some are persuaded. Many were already coming back.
Online video runs the other way. It rarely closes the session, so last-touch almost never credits it, while it can be doing real work earlier that only a holdout can see.
Why this changes decisions, not just reporting
If the gap were a constant multiple, it would be an accounting inconvenience. Every channel would look proportionally worse, the ranking would survive, and since budgets are set by ranking, nothing would change.
The gap is not constant. Ranking by reported CAC and again by incremental CAC produces different orders:
| Channel | By reported CAC | By incremental CAC | Change |
|---|---|---|---|
| Branded search | #2 | #5 | ↓ 3 worse |
| Retargeting | #3 | #6 | ↓ 3 worse |
| Email lifecycle | #1 | #1 | No change |
| Non-brand search | #5 | #4 | ↑ 1 better |
| Paid social prospecting | #4 | #3 | ↑ 1 better |
| Online video | #6 | #2 | ↑ 4 better |
Branded search moves from second-best to fifth. Online video moves from worst to second-best. Retargeting moves from third to last.
A team optimizing against reported CAC will move money toward branded search and retargeting and away from video. That is backwards, and it is the default behavior of most bidding systems and most budget reviews. It happens in small increments, without anyone making an obviously bad call.
This is the mechanism behind the meeting at the top of this page. Reported ROAS improves while incremental revenue stays flat, because budget keeps migrating toward the channels best positioned to claim credit for demand that already existed.
What attribution is actually good for
None of this makes attribution useless, and treating it that way is its own failure mode. It answers a real question, at a granularity and frequency no experiment can match.
Attribution is well suited to:
- Operational diagnosis. Conversions dropped on Tuesday; which source dropped? Attribution answers that in minutes.
- Journey description. How many touchpoints precede a purchase, and over what timespan? That shapes how long a holdout needs to run.
- Relative movement within a channel. Creative A versus creative B, same funnel position, same credit rule. The bias mostly cancels.
- Delivery optimization. Ad platforms need a dense, fast signal. Attributed conversions are the only thing available at that latency.
What it cannot answer is “should we spend more on this.” That question is about a counterfactual, and attribution has no way to construct one.
What to do instead
The practical answer is not to replace attribution but to calibrate it.
- Run holdouts where the stakes are highest. You cannot test everything continuously. Test the channels with the most spend and the most suspicion: usually branded search, retargeting, and anything whose reported performance looks too good.
- Compute an incrementality ratio per channel. Measured incremental conversions divided by last-touch credited conversions, over the same window.
- Apply the ratio to the attributed numbers between tests. Now your daily reporting is on a scale that approximates incremental value, and you keep attribution’s speed and granularity.
- Re-test on a cadence. The ratio drifts as spend, creative, and competitive conditions change. A ratio measured eighteen months ago is a guess.
That gives you a working system rather than a position to defend: fast attributed numbers for daily operation, corrected periodically by slow causal numbers for budget decisions.
Context
Best suited for
Any program spending meaningfully on channels positioned near the conversion: branded search, retargeting, lifecycle email.
Data requirements
- Last-touch or platform-reported conversions by channel
- At least one channel you can withhold cleanly, by geography or by user
- A conversion window long enough to cover most of the journey
What changes by context
- B2C ecommerce
- Short journeys make holdouts fast and cheap, but brand search contamination across market borders is a real threat. Expect the largest gaps on branded search and retargeting.
- B2B SaaS
- The unit is usually an account or an opportunity, not a user, and the lag between touch and revenue can exceed a quarter. Attribution here is often less misleading about direction and far more misleading about timing.
- Sales-led
- A rep is in the path, so last-touch frequently credits whichever marketing touch preceded a conversation the rep actually drove. The incrementality question shifts from "did the ad work" to "which function moved this transition."
- B2C subscription
- Credited conversions include people who would have subscribed later anyway. Pull-forward looks like lift in a short window and disappears in a long one, so the test window has to outlast the natural purchase cycle.
When it breaks down
- Channels that cannot be withheld at any level. Some brand and sponsorship spend cannot be tested.
- Very low volume, where the holdout's confidence interval is wider than the decision you are trying to make.
- Programs where every channel changed at once, leaving no stable period to calibrate against.
What this does not establish
- That the incrementality ratio for a channel is stable. It moves with spend level, creative, seasonality, and competitor behavior.
- That a low-incrementality channel should be cut to zero. It establishes that the marginal dollar is likely worth less than reported, not that the channel has no role.
- That the holdout estimates themselves are correct. Each carries its own assumptions and its own interval, and those intervals are wide.
- Anything about channels that were not tested. A ratio measured on branded search says nothing about display.
What to test next
If this is new territory, the highest-value first test is almost always branded search, because it usually carries large spend, reports excellent performance, and has the widest plausible range of true incrementality. It is also among the easiest to hold out geographically.
The worked example on designing a credible geo holdout carries that kind of test end to end: the design, the power calculation, the diagnostics, the estimate with its interval, and the recommendation that came out of it.
Where these numbers came from Code
The channel table is simulated for teaching, not measured from a real
company. The generating script is committed at
scripts/attribution_gap.py, and the paid social row matches the
holdout estimate produced by scripts/geo_holdout.py so the two
pieces on this site describe the same campaign.
What is not invented is the ordering: the pattern of large gaps at the bottom of the funnel and small or negative gaps at the top is one of the most consistently reproduced findings in the incrementality literature.