Skip to main content

Introduction to Incrementality (Beta)

What incrementality testing is, when it helps, how a geo holdout test works and how it fits alongside MTA and UMM.

Written by Tais Arslan

What is incrementality testing?

Incrementality testing measures the true impact of a campaign or channel. Specifically, it measures the value that activity adds beyond what would have happened anyway, and answers one simple question: would your results have been different if that marketing activity had not occurred?

There are several ways to measure incrementality. Billy Grace uses a geo holdout to measure incrementality. We pause a channel in one region, the holdout, and keep it running everywhere else. For 30 days, we compare what happens in the holdout against what we would have expected if nothing had changed. That difference is the incremental effect.

This is not a user-level A/B test. Everyone in the holdout sees the same experience for the test period, which keeps the comparison clean and the setup straightforward.


When incrementality is useful

Incrementality testing is most helpful when you are trying to separate what your marketing caused from what would have happened anyway. It comes into its own with questions like these.

Are my campaigns driving results?

Problem: The sales, clicks and sign-ups you see might be down to the campaign, or they might have come in organically without it.

Insight: Incrementality shows the causal impact of the campaign, so you can see whether spend is creating new demand or mostly capturing demand that was already there.

Which channels deliver the most value?

Problem: With several channels live at once, it is not always clear which one is truly adding incremental growth.

Insight: By isolating the impact of one channel at a time, you can compare which ones add net-new outcomes and adjust budget with more confidence.

What is the optimal budget allocation?

Problem: It is hard to know how to split budget across channels for the best return.

Insight: tests can show where extra spend still moves the needle and where a channel may already be saturated, so you can rebalance and tap into underfunded areas with incremental potential.


Lower-funnel vs upper-funnel

The way we run an incrementality test is the same for all channels: we pause ads in one region and compare performance with the rest of the country. The main difference between lower-funnel and upper-funnel channels is what we measure and what kind of impact you should expect. Google and Meta are a clear example.

  • For Google, campaigns like Search, Shopping and Performance Max sit closer to the moment of conversion. That means the impact of switching ads on or off is usually visible in conversions. We also look at organic traffic to see whether paid ads are simply replacing organic clicks.

  • For Meta, campaigns typically focus on awareness and consideration earlier in the customer journey. Conversions often happen later, outside the test window. So we mainly measure website traffic (sessions) to understand whether Meta is driving additional demand. Conversions are still monitored, but they are not the primary success metric.

In short, Google tests focus on incremental conversions and Meta tests focus on incremental traffic and demand. The testing approach itself is identical for both.


How we set up and run a test

You handle the geo exclusions and the pause in your ad platforms. We handle the rest, from feasibility through to interpretation. A test runs in five steps.

  1. Feasibility check. Before recommending a holdout, we run simulations on your historical data to find suitable regions and channels and to confirm a 30-day test can detect a real effect. This is also where we check that the channel under consideration is large enough in that region for the test to be meaningful.

    Example: if Google Branded Search is 7% of your last-30-day conversions, feasibility might show Noord-Holland needs a 10% drop to detect an effect while Zuid-Holland needs only 5%. Zuid-Holland is then the better region to test.

  2. Run the experiment. You turn off or geo-exclude advertising in the selected region for 30 days. It helps to pick a period without unusual swings in sales or spend, and to avoid sale periods, since regions can react to the same promotion differently and blur the comparison.

  3. Monitoring. While the test runs, we watch spend and impressions for anything unexpected, such as a spike, a drop or a change on a cluster you did not intend to pause. If something looks off, we flag it, and we may suggest ending the test early or extending it so the comparison stays trustworthy.

  4. Estimation. After the holdout, we estimate what would have happened if ads had stayed on, using a synthetic control built from regions that tracked the holdout closely beforehand. The gap between expected and actual is the incremental effect, reported as an effect size with a confidence interval.

  5. Reading the results. We evaluate the effect, its confidence interval and statistical significance to see whether the channel created a measurable incremental impact. How we use it then depends on what you tested: for conversion-focused holdouts, whether attributed conversions were truly incremental or mostly shifting between channels; for upper-funnel holdouts, a read on sessions and cross-channel demand that can recalibrate how UMM models the effect of impressions for the next cycle.


Designing a test that gives a clear answer

The basic idea is simple: pause in one region and compare to the rest of the country. Good planning is what makes the difference. Even with a sound design, the wrong region, the wrong timing or changes mid-test can leave you with a result you cannot act on. A few common issues are worth knowing upfront.

  • Region or channel too small. The pause may not produce a change large enough to detect. The feasibility check is there to catch this before you turn any ads off.

  • Busy periods like sales and promotions. Regions can react differently to the same event, so we recommend running tests in more stable periods.

  • Changes during the test. New campaigns, promotions or spend shifts can blur the comparison. This is what the monitoring step is watching for.

  • Wrong metric for the channel. Measuring Meta only on in-window conversions, for example, misses the demand it builds. We align the primary metric with the channel so the test is measured on what it actually drives.

Feasibility looks back, not forward

The feasibility check is built on historical data. It tells us whether a channel or region has behaved consistently enough in the past to produce a reliable result, it can't tell us what's about to change.

That matters because a test only measures what it's designed to isolate. If something else shifts at the same time, a broader overhaul of your campaigns, a new website, a big change in creative or targeting, we can no longer tell whether a change in performance came from the test or from that other change. The test doesn't fail loudly; it just becomes harder to trust.

Before you launch, it's worth asking: is anything else changing for this channel, region, or website in the next month? If the answer is yes, the test will still run, but the result may not be one you can act on. In that case, it's usually better to wait until things have settled, or to test a channel or region that isn't affected by the upcoming change.


An example of an incrementality test

Say you are running Meta ads and want to measure their incremental value. You decide to run a geo holdout test.

  • Holdout region: in Noord-Holland, you pause all Meta ads for the duration of the test, or simply exclude Noord-Holland from targeting.

  • Control region: in the rest of the country, you keep running Meta ads as usual.

For the 30-day test, Billy Grace tracks and compares behaviour in both regions. By comparing the two, we can work out how much of the website traffic in Noord-Holland was lost with Meta ads paused, while the rest of the country serves as the baseline where ads kept running.


How incrementality fits with MTA and UMM

Incrementality does not replace your attribution. It checks it against the real world. MTA shows which clicks in a journey mattered, UMM adds the impact of impressions on top, and both are built from your data and the platforms. A geo holdout brings in something neither can on its own: evidence from a real pause.

That makes incrementality a causal measurement you can set against what MTA and UMM already show. When a test confirms the demand effects UMM has modelled, those findings feed back to recalibrate UMM for the next cycle. The three work together, each one sharpening the others.

Talk to your Billy Grace contact about where a test could sharpen your next budget decision.

Did this answer your question?