Once your Incrementality test has finished, the results page shows how much of an effect the tested channel actually had, separate from what would have happened anyway. This article walks through each section of the results page and how to read it.
For how to set up and launch a test, see [the incrementality testing setup article].
Test Overview
At the top of the page you'll find your test settings and the overall incremental lift across all sources.
This panel shows:
Channel: the channel being tested
Cluster: the specific ad cluster included in the test (e.g. Branded Search), or "All" if the full channel was tested
Start date / End date: the period the test ran
Country / Test region: where the test was run (holdout)
Goal: the metric the test is measuring (e.g. Sessions)
Incremental lift: the overall effect the channel had on the goal, across the full test period
Confidence level: how much weight to put on the lift figure. See Confidence level below
Incremental Lift
This chart shows how the incremental lift evolved day by day over the test period.
There are two dropdowns:
Source: filter to see the tested channel's effect on a specific source (e.g. whether pausing branded search also changed sessions coming from organic search), or leave on "All" for the combined effect
Goal: switch between goals (e.g. Sessions, Purchases) to see the incremental lift calculated for each
Changing either dropdown recalculates the incremental lift shown, since you're now looking at a different slice of the data.
Reading the line and the band
The line is the estimated incremental lift for that day. The shaded band around it is the confidence interval: the range the true effect is likely to fall within. A wider band means less certainty; a narrower band means more.
Depending on the result, this can look a few different ways.
Take a brand running an incrementality test on branded search.
Positive lift: the line sits above zero. The tested channel is driving incremental results: without it, you'd be missing out on this volume.
For example, pausing branded search causes purchases to drop, and no other source picks up the difference, proof that branded search was bringing in demand nothing else was capturing.No lift: the line sits at or close to zero. This doesn't mean the channel had no value. It usually means another source, such as organic search, stepped in and picked up the demand the tested channel would otherwise have driven.
For example, pausing branded search doesn't move total purchases at all, because organic search sessions rise by roughly the same amount, those customers still found their way to the brand, just through a different route.Negative lift: the line sits below zero. The tested channel was actively working against another source, for example by substituting traffic that would have converted anyway.
For example, pausing branded search causes organic search sessions to rise.
Confidence level
Every incrementality test result comes with a confidence level. It tells you how much weight to put on the number in front of you, not whether the result is good or bad news.
Here's why it exists. A geo holdout test compares what actually happened in the test region with what would have happened there if the channel had kept running. That second part is an estimate, built from the control region where the channel stayed live. Any estimate carries a margin of uncertainty, and the confidence level is that margin translated into plain language.
There are three levels.
Highly confident: Both the direction of the effect and its size are well established. The test shows a clear lift, or a clear drop.
How to use it: Act on the full result. The direction and the number are both solid enough to plan budget against.Moderately confident: The direction is well evidenced. The channel is moving the outcome and you can tell which way. The size is less settled. This shows up most often when the test window is short, the spend level is modest, or the outcome you're measuring is naturally volatile.
How to use it: Trust the direction and treat the number as an order of magnitude rather than a precise figure. It's enough to justify keeping a channel live or shifting budget towards it, but not enough to set a precise iROAS target. A longer or higher-spend follow-up test will tighten the estimate.No incremental lift: The test found no strong evidence of an incremental effect. The true effect may be genuinely close to zero, or it may be real but too small to separate from normal day-to-day variation.
How to use it: Read it as no measurable lift at this spend level, in this window, and use that to inform where the next euro goes. It's a finding rather than a gap in the data.
Sessions
This chart shows the raw goal metric recorded for the test region and the control region across the test period, plus some time before and after for context.
Test region: the region where the channel was paused/active for the test
Control region: a comparable region left unchanged, used as a baseline
The incremental lift is calculated from the gap between these two lines: the more the test region diverges from the control region during the test period, the larger the incremental lift.
Spend
This chart shows spend on the tested channel/ campaigns over time. You'll typically see it drop to zero at the start date and stay there for the duration of the test.
This is expected: pausing spend in the test region is what makes it possible to isolate the channel's effect. Use this chart as a control check: if spend continues in the test region, or changes in the control region, that would affect how reliable the result is.
Status test result
Every incrementality test has a status that reflects where it is in the process.
Pending: The test has not started yet. Spend has not yet been paused in the holdout region; this action is required on your end before it can begin.
In progress: The test is running.
Validation: The test period is over and our data team is validating the results. You can switch your spend back on.
Completed: The test has finished successfully and the results are available.
Failed: The test ended automatically because contamination was detected, usually spend appearing in the holdout region when it should have been paused. No results are available for failed tests.
Stopped: The test was ended manually by a user before it finished. No results are available for stopped tests.
Test examples
Example 1: a branded search test that came back with no incremental lift in conversions
The setup
Paid branded search was paused in a set of holdout regions while it kept running everywhere else. The outcome measured was conversions.
The result
We detected no incremental lift in conversions. In the same test, organic sessions to the website rose, and that lift was clear enough to be reported with high confidence.
How to read it
Branded search captures people who are already looking for you by name. When the paid ad isn't there, a large share of that demand still arrives anyway. It just comes through the organic result instead. The demand didn't disappear; it changed route.
That's why the two findings sit together so neatly. No incremental lift in conversions and the significant organic lift are describing the same behaviour from two angles, which is a good sign that the test measured what it set out to measure.
What to do with it
Treat it as a question about allocation, not as an instruction to switch branded search off. The right level of branded coverage depends on competitor bidding, your organic visibility and how much of your brand demand you're willing to leave uncontested.
Use it as a reason to test at a different spend level rather than a different channel. Reducing branded spend and re-testing tells you more than pausing entirely.
Re-test periodically. Competitor activity on your brand terms is the variable most likely to change this result.
Example 2: an upper-funnel social test measured on session signals
The setup
An upper-funnel social campaign was paused in a set of holdout regions. Rather than anchoring on conversions, the test measured short-term session signals across the four channel groupings we report on by default: Google, direct, organic and other.
Why the outcome is different
Upper-funnel campaigns rarely close a sale inside a short test window. The path from first exposure to purchase is long and indirect, so a conversion-based test on a two to four week window is often measuring a signal too faint to detect. Anchoring on conversions here would produce a result that reflects the test design rather than the campaign.
Session signals move faster. They show whether the campaign is creating demand that other parts of the mix then pick up.
The result
Each grouping is compared with its own counterfactual. Where a grouping comes in below that counterfactual, that's evidence the campaign was contributing to it, and that grouping gets its own confidence level. Highly confident when both the direction and the size of the gap are well established, moderately confident when the direction is clear but the size sits in a wider range.
Two of the groupings are usually where the evidence shows up first in an upper-funnel test:
A fall in Google sessions points to fewer people searching for the brand by name. Branded search sits inside this grouping, and coverage in the holdout regions stays unchanged during the test, so a gap here reflects fewer people looking rather than less visibility.
A fall in direct sessions points to fewer people arriving by typing the URL, using a bookmark, or opening the app. Direct is the closest thing to a clean read on brand recall in a short window.
Organic and other are measured and reported the same way. Which groupings move, and by how much, depends on the brand, the market and the length of the test. A flat result on one grouping doesn't undercut a clear result on another.
How to read it
A result like this means the campaign is creating demand rather than closing it. People exposed to it go on to search for the brand by name or come to the site directly, so part of what shows up as Google and direct volume in your reporting started here.
The comparison with the branded search example above is what makes that reading hold. There, one grouping fell and another rose to meet it, so the demand rerouted rather than disappeared. In an upper-funnel test, a fall with no offsetting rise anywhere else in the mix points to demand that wouldn't have existed without the campaign.
What to do with it
Judge upper-funnel campaigns on the outcome they actually drive in a short window, and keep conversion-based measurement for the campaigns that sit closer to purchase.
Read your Google, direct and other session volume with this in mind. Some of it is downstream of upper-funnel activity, and a cut there will show up in those lines a week or two later.
FAQ
What does "Other sources" mean in the Source dropdown?
"Other sources" groups together smaller sources that aren't worth analysing individually. The dropdown shows each major source on its own, with everything else combined into "Others."




