Designing a Credible Geo Holdout
A paid social incrementality test carried end to end: market pairing, power, pre-trend diagnostics, a difference-in-differences estimate with its interval, incremental CAC, and what the design cannot settle.
- The decision
- Should paid social prospecting keep its budget, grow, or shrink?
- The method
- Matched-pair geo holdout, two-way fixed effects difference-in-differences on log conversions.
STEP 01
The business decision
Paid social prospecting is running at roughly $47,500 per week. The channel reports a solid return on last-touch. Finance has asked whether the budget should grow, hold, or move to search.
The decision is not whether paid social works. It is whether the next dollar here buys more than the same dollar somewhere else. That needs the number of conversions this spend creates, not the number it gets credited with.
STEP 02
The tempting metric
The channel reports 4,820 last-touch conversions for the quarter against $285,000 of spend. That is a reported CAC of about $59, comfortably inside target. On this number, the answer is obviously to spend more.
STEP 03
Why that may be wrong
Reported CAC counts conversions the platform could claim. That includes people who saw an ad and would have bought anyway, people already mid-journey, and people who converted through another route but had an impression in the lookback window.
None of that is dishonest reporting. Credit is not causation, and the size of the gap stays unknown until you measure it. Prospecting reaches colder audiences, so its gap is usually smaller than retargeting or branded search. “Usually smaller” is not a number you can put in a budget model.
STEP 04
Data available and missing
Available. Conversions attributable to a market by billing postcode, daily, back three years. Spend controllable at the market level in the ad platform. Clean market definitions that match how the media is bought.
Missing. Any user-level link between ad exposure and purchase. Consent rates make that unreliable, which is what pushes this toward a geo design rather than a user-level one. Also missing: any prior experiment on this channel, so there is no existing estimate to anchor on.
STEP 05
The target quantity
The estimand is the average effect of running paid social prospecting on conversions in a market, during the weeks it runs, for markets like these.
That phrasing carries three limits, and each one matters later:
- During the weeks it runs. Not a persistent effect, not a brand effect that builds over a year.
- At this spend level. Effects are not linear in spend, so this says nothing about doubling.
- For markets like these. Forty markets chosen for testability, not a random sample of geography.
STEP 06
The strongest feasible design
40 markets, split 19 treated and 21 held out, with an 8-week pre-period and a 6-week test.
Markets are assigned in matched pairs: rank by baseline conversion volume, then randomize within each adjacent pair. Pure random assignment across markets that differ by an order of magnitude leaves too much to chance, since one large market landing in a single arm can dominate the comparison. Pairing forces balance on the variable most likely to matter. Randomizing inside the pair keeps the randomization that makes the comparison credible.
The resulting arms are balanced on baseline volume to within 17.9%.
Was it powered?
Before running, the question is what lift this design could detect at all. Using the observed within-market week-to-week variability (σ = 0.095 on the log scale) and 20 markets per arm:
Show the data
| Period | MDE |
|---|---|
| 2w | 6.1% |
| 3w | 5.0% |
| 4w | 4.3% |
| 5w | 3.8% |
| 6w | 3.5% |
| 7w | 3.2% |
| 8w | 3.0% |
| 9w | 2.8% |
| 10w | 2.7% |
| 11w | 2.6% |
| 12w | 2.5% |
Six weeks reaches an MDE below the roughly 5% lift that would change the budget decision, at a cost of about $285,000 in held-out spend. Going longer buys precision slowly: standard error falls with the square root of duration, so doubling the test narrows the interval by only about 30%.
You can run this calculation for your own design with the geo holdout designer.
STEP 07
Assumptions and failure modes
STEP 08
Worked example and diagnostics
The pre-trend check
The diagnostic that decides whether the headline number deserves belief. Both arms are indexed to their own pre-period average, because the arms differ in absolute size and raw totals would only show two parallel bands.
- Treated
- Held out
Show the data
| Period | Treated | Held out |
|---|---|---|
| W-8 | 95.9 | 99.6 |
| W-7 | 103.5 | 102.5 |
| W-6 | 105.8 | 108.8 |
| W-5 | 110.8 | 105.5 |
| W-4 | 98.2 | 100.3 |
| W-3 | 96.9 | 98.0 |
| W-2 | 94.2 | 94.9 |
| W-1 | 94.8 | 90.4 |
| W1 | 105.9 | 94.2 |
| W2 | 104.2 | 103.0 |
| W3 | 112.3 | 106.4 |
| W4 | 119.6 | 112.7 |
| W5 | 118.8 | 108.4 |
| W6 | 115.0 | 105.3 |
This is a passing diagnostic, not a perfect one. The pre-period gap of +1.3% per week is small against the noise, but it is not zero, and eight pre-period weeks gives the check limited power to catch a slow divergence. It makes parallel trends plausible. It does not make it proven, and no amount of pre-period data would.
The estimator
Two-way fixed effects on log conversions:
log(conversions[market, week]) = α[market] + τ[week] + β · (treated × post) + εMarket fixed effects absorb persistent size differences. Week fixed effects absorb the shared demand curve: the seasonality and category drift that a naive before-and-after comparison would misread as campaign effect. What remains in β is the differential change in treated markets during the test window.
STEP 09
The estimate, with uncertainty
Effect of paid social prospecting on conversions in treated markets
The interval is wide. It clears zero comfortably, so the channel is doing something, but it spans a range where the budget conclusion differs at each end.
- Incremental conversions
- 3,246
- Incremental CAC
- $88
- Spend held out
- $285k
Show the data
| Estimated lift | Point | 95% CI |
|---|---|---|
| Paid social prospecting | +7.2% | +4.0% to +10.5% |
Converting the percentage lift into conversions: treated markets recorded 48,383 conversions during the test. At an estimated lift of +7.2%, roughly 3,246 of those were incremental, which puts about 45,137 in the “would have happened anyway” column.
Against $285,000 of spend, that is an incremental CAC of about $88, against a reported CAC of $59.
The interval matters more than the point. At the optimistic end (+10.5%) incremental CAC is about $62. At the pessimistic end (+4.0%) it is about $154. Those two numbers support different decisions, and this test does not separate them.
STEP 10
The recommendation
Hold the budget. Do not grow it on this evidence. Re-test at a higher spend level before committing more.
The reasoning:
- The channel works. The interval clears zero by a comfortable margin, so this is not a channel to cut.
- Incremental CAC of about $88 runs well above the reported $59, but stays inside the range the business tolerates for new customers.
- Growing the budget requires the marginal return. This test measures the average effect at current spend. For a channel that may be nearing saturation, those two can differ a great deal.
The summary for the budget meeting: this channel creates value, reported CAC overstates it by roughly a third, and we do not yet know what the next dollar buys.
STEP 11
What this does not establish
What this does not establish
- That a larger budget would produce proportionally more conversions. This measures the effect at current spend. Response curves bend, and this test cannot see the bend.
- That the effect persists after the campaign stops. The estimand is deliberately limited to the weeks the campaign ran; a pull-forward effect and a genuine demand increase look identical inside this window.
- That the same lift applies in markets outside the test. These 40 markets were selected for clean measurement, not sampled at random from all geography.
- That paid social is better or worse than search. Each channel needs its own test; comparing this estimate to another channel's reported CAC would repeat the exact error this test exists to correct.
- That parallel trends held during the test window. The pre-period check makes it plausible. It cannot be verified where it matters.
STEP 12
What to test next
- A spend-level test. Split treated markets into two spend tiers and estimate the response curve instead of a single point. That answers the question this test could not.
- A longer post-window. Continue measuring held-out markets for six weeks after the campaign resumes, to separate genuine demand creation from pull-forward.
- Branded search. Higher spend, better reported performance, and a much wider plausible range of true incrementality. Per dollar of held-out spend, it is almost certainly the better next test.
- A persistent holdout. A permanent 5% held-out set of markets turns incrementality from a project into a standing number, and makes ratio calibration practical.
STEP 13
Technical appendix
The estimator, standard errors, and reproduction Code
Coefficient on the treated × post interaction:
0.0694 on the log scale, standard error
0.0156, which exponentiates to the
+7.2% lift reported above.
Standard errors here are conventional OLS. In production they should be clustered at the market level, because observations within a market are correlated across weeks. Conventional standard errors run optimistic in that setting, so the interval above is narrower than it should be. That caveat cuts against the strength of the finding.
Conversions enter as log(y + 0.5) to keep zero-conversion
market-weeks defined. With market volumes at this scale that offset is
immaterial; at low volume it would not be, and a Poisson or negative binomial
model would be the better choice.
Every number on this page is regenerated from seed
20260823 by a committed script, so the dataset and the
figures cannot drift apart. The source is private; happy to walk through it.
How well did the estimator do? Technical
Because this is simulated, the true effect is known, which no real test ever is. The data were generated with a true lift of +7.5%. The estimate came back at +7.2%, with an interval of +4.0% to +10.5% that contains the truth.
That is a good outcome and a slightly misleading one. The point estimate landing this close is partly luck. The interval is wide because a single six-week test on forty markets cannot pin the effect down further. Had the estimate come back at +10.0% or +5.0%, the test would have been just as well executed and the interval would still have covered the truth. The interval is the finding.
Context
Best suited for
Channels that can be switched off in a geography without breaking the rest of the program, where conversions are attributable to a market and spend is controllable at market level.
Data requirements
- Conversions attributable to a geography: billing postcode, shipping address, or store
- At least 8 weeks of clean pre-period history at market grain
- Spend controllable at the geographic level in the buying platform
- Roughly 20+ markets per arm before the MDE gets larger than effects worth finding
What changes by context
- B2C ecommerce
- The default case, and the easiest. Short purchase cycles mean a 4–6 week test captures most of the effect. Watch for brand search contamination across market borders.
- B2B SaaS
- The outcome should usually be an opportunity or a qualified account rather than a lead, and those lag by a quarter or more. Either extend the window past the lag or use an earlier validated proxy, and say which you chose.
- Marketplace
- Supply-side spillovers break the no-interference assumption in a way geography does not contain: a treated market can pull inventory or sellers away from a held-out one, biasing the estimate upward on both sides.
- Local / omnichannel
- Geo designs are strongest here, since the market is the natural unit. The complication moves to outcome measurement, since offline conversions need a reliable market-level source.
- B2C subscription
- Pull-forward is the main threat. A promotion that accelerates existing intent looks like lift in a six-week window and vanishes in a six-month one, so the test window must outlast the natural purchase cycle.
When it breaks down
- Fewer than roughly 20 markets per arm, where the MDE exceeds any effect that would change a decision.
- National campaigns that cannot be geo-targeted at all, such as broadcast and most sponsorships.
- Markets that share media footprints, where holdout markets receive spillover and the estimate is biased toward zero.
- Any period when something else changed unevenly across markets: a regional promotion, a competitor launch, a distribution change.