Skip to content
Was It Causal?

Designing a Credible Geo Holdout

A paid social incrementality test carried end to end: market pairing, power, pre-trend diagnostics, a difference-in-differences estimate with its interval, incremental CAC, and what the design cannot settle.

Incrementality Applied Worked example B2C ecommerce Marketing-led
The decision
Should paid social prospecting keep its budget, grow, or shrink?
The method
Matched-pair geo holdout, two-way fixed effects difference-in-differences on log conversions.

STEP 01

The business decision

Paid social prospecting is running at roughly $47,500 per week. The channel reports a solid return on last-touch. Finance has asked whether the budget should grow, hold, or move to search.

The decision is not whether paid social works. It is whether the next dollar here buys more than the same dollar somewhere else. That needs the number of conversions this spend creates, not the number it gets credited with.

STEP 02

The tempting metric

The channel reports 4,820 last-touch conversions for the quarter against $285,000 of spend. That is a reported CAC of about $59, comfortably inside target. On this number, the answer is obviously to spend more.

STEP 03

Why that may be wrong

Reported CAC counts conversions the platform could claim. That includes people who saw an ad and would have bought anyway, people already mid-journey, and people who converted through another route but had an impression in the lookback window.

None of that is dishonest reporting. Credit is not causation, and the size of the gap stays unknown until you measure it. Prospecting reaches colder audiences, so its gap is usually smaller than retargeting or branded search. “Usually smaller” is not a number you can put in a budget model.

STEP 04

Data available and missing

Available. Conversions attributable to a market by billing postcode, daily, back three years. Spend controllable at the market level in the ad platform. Clean market definitions that match how the media is bought.

Missing. Any user-level link between ad exposure and purchase. Consent rates make that unreliable, which is what pushes this toward a geo design rather than a user-level one. Also missing: any prior experiment on this channel, so there is no existing estimate to anchor on.

STEP 05

The target quantity

The estimand is the average effect of running paid social prospecting on conversions in a market, during the weeks it runs, for markets like these.

That phrasing carries three limits, and each one matters later:

  • During the weeks it runs. Not a persistent effect, not a brand effect that builds over a year.
  • At this spend level. Effects are not linear in spend, so this says nothing about doubling.
  • For markets like these. Forty markets chosen for testability, not a random sample of geography.

STEP 06

The strongest feasible design

40 markets, split 19 treated and 21 held out, with an 8-week pre-period and a 6-week test.

Markets are assigned in matched pairs: rank by baseline conversion volume, then randomize within each adjacent pair. Pure random assignment across markets that differ by an order of magnitude leaves too much to chance, since one large market landing in a single arm can dominate the comparison. Pairing forces balance on the variable most likely to matter. Randomizing inside the pair keeps the randomization that makes the comparison credible.

The resulting arms are balanced on baseline volume to within 17.9%.

Was it powered?

Before running, the question is what lift this design could detect at all. Using the observed within-market week-to-week variability (σ = 0.095 on the log scale) and 20 markets per arm:

2.0 2.5 3.0 3.5 4.0 4.5 5.0 5.5 6.0 6.5 2w 3w 4w 5w 6w 7w 8w 9w 10w 11w 12w MDE MDE, 2w: 6.1% MDE, 3w: 5.0% MDE, 4w: 4.3% MDE, 5w: 3.8% MDE, 6w: 3.5% MDE, 7w: 3.2% MDE, 8w: 3.0% MDE, 9w: 2.8% MDE, 10w: 2.7% MDE, 11w: 2.6% MDE, 12w: 2.5% MDE
Show the data
Period MDE
2w 6.1%
3w 5.0%
4w 4.3%
5w 3.8%
6w 3.5%
7w 3.2%
8w 3.0%
9w 2.8%
10w 2.7%
11w 2.6%
12w 2.5%
Minimum detectable effect falls with duration, but the returns flatten. Six weeks reaches roughly 3.5%, and twelve would only reach 2.5%. Simulated data · scripts/geo_holdout.py

Six weeks reaches an MDE below the roughly 5% lift that would change the budget decision, at a cost of about $285,000 in held-out spend. Going longer buys precision slowly: standard error falls with the square root of duration, so doubling the test narrows the interval by only about 30%.

You can run this calculation for your own design with the geo holdout designer.

STEP 07

Assumptions and failure modes

STEP 08

Worked example and diagnostics

The pre-trend check

The diagnostic that decides whether the headline number deserves belief. Both arms are indexed to their own pre-period average, because the arms differ in absolute size and raw totals would only show two parallel bands.

  • Treated
  • Held out
Test period 90 95 100 105 110 115 120 W-8 W-6 W-4 W-2 W1 W3 W5 Index Campaign starts Treated, W-8: 95.9 Treated, W-7: 103.5 Treated, W-6: 105.8 Treated, W-5: 110.8 Treated, W-4: 98.2 Treated, W-3: 96.9 Treated, W-2: 94.2 Treated, W-1: 94.8 Treated, W1: 105.9 Treated, W2: 104.2 Treated, W3: 112.3 Treated, W4: 119.6 Treated, W5: 118.8 Treated, W6: 115.0 Treated Held out, W-8: 99.6 Held out, W-7: 102.5 Held out, W-6: 108.8 Held out, W-5: 105.5 Held out, W-4: 100.3 Held out, W-3: 98.0 Held out, W-2: 94.9 Held out, W-1: 90.4 Held out, W1: 94.2 Held out, W2: 103.0 Held out, W3: 106.4 Held out, W4: 112.7 Held out, W5: 108.4 Held out, W6: 105.3 Held out
Show the data
Period Treated Held out
W-8 95.9 99.6
W-7 103.5 102.5
W-6 105.8 108.8
W-5 110.8 105.5
W-4 98.2 100.3
W-3 96.9 98.0
W-2 94.2 94.9
W-1 94.8 90.4
W1 105.9 94.2
W2 104.2 103.0
W3 112.3 106.4
W4 119.6 112.7
W5 118.8 108.4
W6 115.0 105.3
The arms track each other closely through the pre-period, then separate once the campaign starts. Mean weekly growth-rate gap before the test was +1.3%, well inside the 5.2pp week-to-week noise. Simulated data · scripts/geo_holdout.py

This is a passing diagnostic, not a perfect one. The pre-period gap of +1.3% per week is small against the noise, but it is not zero, and eight pre-period weeks gives the check limited power to catch a slow divergence. It makes parallel trends plausible. It does not make it proven, and no amount of pre-period data would.

The estimator

Two-way fixed effects on log conversions:

log(conversions[market, week]) = α[market] + τ[week] + β · (treated × post) + ε

Market fixed effects absorb persistent size differences. Week fixed effects absorb the shared demand curve: the seasonality and category drift that a naive before-and-after comparison would misread as campaign effect. What remains in β is the differential change in treated markets during the test window.

STEP 09

The estimate, with uncertainty

Effect of paid social prospecting on conversions in treated markets

+7.2% 95% CI +4.0% to +10.5%

The interval is wide. It clears zero comfortably, so the channel is doing something, but it spans a range where the budget conclusion differs at each end.

Incremental conversions
3,246
Incremental CAC
$88
Spend held out
$285k
0 % 2 % 4 % 6 % 8 % 10 % 12 % No effect Paid social prospecting Paid social prospecting: +7.2% (95% CI +4.0% to +10.5%) +7.2% (+4.0% to +10.5%)
Show the data
Estimated lift Point 95% CI
Paid social prospecting +7.2% +4.0% to +10.5%
Read the interval, not the point. The estimate clears zero, which settles direction but not magnitude. Simulated data · scripts/geo_holdout.py

Converting the percentage lift into conversions: treated markets recorded 48,383 conversions during the test. At an estimated lift of +7.2%, roughly 3,246 of those were incremental, which puts about 45,137 in the “would have happened anyway” column.

Against $285,000 of spend, that is an incremental CAC of about $88, against a reported CAC of $59.

The interval matters more than the point. At the optimistic end (+10.5%) incremental CAC is about $62. At the pessimistic end (+4.0%) it is about $154. Those two numbers support different decisions, and this test does not separate them.

STEP 10

The recommendation

Hold the budget. Do not grow it on this evidence. Re-test at a higher spend level before committing more.

The reasoning:

  • The channel works. The interval clears zero by a comfortable margin, so this is not a channel to cut.
  • Incremental CAC of about $88 runs well above the reported $59, but stays inside the range the business tolerates for new customers.
  • Growing the budget requires the marginal return. This test measures the average effect at current spend. For a channel that may be nearing saturation, those two can differ a great deal.

The summary for the budget meeting: this channel creates value, reported CAC overstates it by roughly a third, and we do not yet know what the next dollar buys.

STEP 11

What this does not establish

What this does not establish

  • That a larger budget would produce proportionally more conversions. This measures the effect at current spend. Response curves bend, and this test cannot see the bend.
  • That the effect persists after the campaign stops. The estimand is deliberately limited to the weeks the campaign ran; a pull-forward effect and a genuine demand increase look identical inside this window.
  • That the same lift applies in markets outside the test. These 40 markets were selected for clean measurement, not sampled at random from all geography.
  • That paid social is better or worse than search. Each channel needs its own test; comparing this estimate to another channel's reported CAC would repeat the exact error this test exists to correct.
  • That parallel trends held during the test window. The pre-period check makes it plausible. It cannot be verified where it matters.

STEP 12

What to test next

  1. A spend-level test. Split treated markets into two spend tiers and estimate the response curve instead of a single point. That answers the question this test could not.
  2. A longer post-window. Continue measuring held-out markets for six weeks after the campaign resumes, to separate genuine demand creation from pull-forward.
  3. Branded search. Higher spend, better reported performance, and a much wider plausible range of true incrementality. Per dollar of held-out spend, it is almost certainly the better next test.
  4. A persistent holdout. A permanent 5% held-out set of markets turns incrementality from a project into a standing number, and makes ratio calibration practical.

STEP 13

Technical appendix

The estimator, standard errors, and reproduction Code

Coefficient on the treated × post interaction: 0.0694 on the log scale, standard error 0.0156, which exponentiates to the +7.2% lift reported above.

Standard errors here are conventional OLS. In production they should be clustered at the market level, because observations within a market are correlated across weeks. Conventional standard errors run optimistic in that setting, so the interval above is narrower than it should be. That caveat cuts against the strength of the finding.

Conversions enter as log(y + 0.5) to keep zero-conversion market-weeks defined. With market volumes at this scale that offset is immaterial; at low volume it would not be, and a Poisson or negative binomial model would be the better choice.

Every number on this page is regenerated from seed 20260823 by a committed script, so the dataset and the figures cannot drift apart. The source is private; happy to walk through it.

How well did the estimator do? Technical

Because this is simulated, the true effect is known, which no real test ever is. The data were generated with a true lift of +7.5%. The estimate came back at +7.2%, with an interval of +4.0% to +10.5% that contains the truth.

That is a good outcome and a slightly misleading one. The point estimate landing this close is partly luck. The interval is wide because a single six-week test on forty markets cannot pin the effect down further. Had the estimate come back at +10.0% or +5.0%, the test would have been just as well executed and the interval would still have covered the truth. The interval is the finding.

Context

Best suited for

Channels that can be switched off in a geography without breaking the rest of the program, where conversions are attributable to a market and spend is controllable at market level.

Data requirements

  • Conversions attributable to a geography: billing postcode, shipping address, or store
  • At least 8 weeks of clean pre-period history at market grain
  • Spend controllable at the geographic level in the buying platform
  • Roughly 20+ markets per arm before the MDE gets larger than effects worth finding

What changes by context

B2C ecommerce
The default case, and the easiest. Short purchase cycles mean a 4–6 week test captures most of the effect. Watch for brand search contamination across market borders.
B2B SaaS
The outcome should usually be an opportunity or a qualified account rather than a lead, and those lag by a quarter or more. Either extend the window past the lag or use an earlier validated proxy, and say which you chose.
Marketplace
Supply-side spillovers break the no-interference assumption in a way geography does not contain: a treated market can pull inventory or sellers away from a held-out one, biasing the estimate upward on both sides.
Local / omnichannel
Geo designs are strongest here, since the market is the natural unit. The complication moves to outcome measurement, since offline conversions need a reliable market-level source.
B2C subscription
Pull-forward is the main threat. A promotion that accelerates existing intent looks like lift in a six-week window and vanishes in a six-month one, so the test window must outlast the natural purchase cycle.

When it breaks down

  • Fewer than roughly 20 markets per arm, where the MDE exceeds any effect that would change a decision.
  • National campaigns that cannot be geo-targeted at all, such as broadcast and most sponsorships.
  • Markets that share media footprints, where holdout markets receive spillover and the estimate is biased toward zero.
  • Any period when something else changed unevenly across markets: a regional promotion, a competitor launch, a distribution change.