Proving the Value of Paid Search With Incrementality Testing
Marketers have wrestled with attribution questions since before digital tracking existed. Like all other marketing tactics, Paid Search hasn’t solved it. Enter incrementality testing.
Incrementality testing may sound intimidating to those of us without a background in statistics. Many think it’s “just the next iteration of attribution modeling,” but there’s more to it. It’s the task of answering a simple question – How much of this would’ve happened anyway? This is called counterfactuality. After all, attribution modeling is almost always skewed by platform bias, such as letting Google Analytics gauge the efficacy of Google Ads campaigns.
Meanwhile, other platforms are also claiming credit for the very same leads and conversions. Every platform claims the win. Even within paid search, Brand Search often maintains a place in the spotlight without appropriately crediting the top-of-funnel tactics that helped to drive those conversions. Incrementality testing eliminates those biases and creates real-world testing environments intended to give credit where it’s actually due.
What Is Incrementality Testing?
Incrementality testing is the practice of gauging whether specific marketing actions are actually impactful. An experiment is created by establishing a baseline for standard performance and adjusting some variable in the marketing mix to observe whether the change to the selected variable affects overall outcomes. Rather than evaluating past campaigns and overall performance to attribute results, incrementality testing is a deliberate act of changing inputs for the sake of measuring what happens.
Paid Search programs are no exception to the need for a deeper understanding of tactical incrementality. Brand Search is easy to pick on, since it tends to show the strongest KPIs. But where do those purchases, leads, and other conversion types actually come from? What role do tactics like Non-Brand Search and Demand Gen play in replenishing the funnel and ultimately informing and persuading users who later convert via branded search terms? The smartest Paid Search managers are using incrementality testing to understand the contributions of each part of their program.
Evaluating the Incrementality of Paid Search Tactics
Setting up an incrementality test may seem daunting at first, but an ongoing testing structure helps eliminate wasted spend and build long-term success in Paid Search programs. There are several factors to consider when developing this type of testing framework.
Define Success
To begin, an advertiser must first determine exactly what they hope to evaluate. A test must define what’s being observed, such as the true contribution of Non-Brand Search or the unattributed value of Video campaigns. And when deciding what to measure, it’s important to identify not only relevant KPIs, but also to determine a source of truth that all stakeholders can agree on. An example of what this might look like for an ecommerce company is “we are measuring the impact of Non-Brand Search by monitoring for an overall lift in Sales and Revenue within Shopify.”
Pro-Tip: To effectively measure lift in any form, an advertiser must have a reasonable understanding of what baseline performance is. This is used to estimate what performance would look like without the influence of whatever is being tested, so that the difference between this estimated baseline and what actually happens can be attributed to whatever is being tested as incremental lift.
H3: Define Test Parameters
Once it’s been decided what will be measured, the next step is selecting how to structure the test to create the clearest view. Geo Holdout and Lift Tests are two of the most popular testing structures among Paid Search managers.
H4: Geo Holdout
Geo holdout tests are a relatively low-lift incrementality testing option for Paid Search managers. The simplest explanation is that certain campaigns or campaign types are paused in specific geos for a set period. After that period, results from the test group are compared to those of a designated control group. The main obstacle here is ensuring that it’s possible to get state-level conversion data from the chosen source of truth.
H4: Lift Tests
Sometimes referred to as user holdout tests, Lift tests are conducted by exposing a set group of users to a set of ads and comparing results among those users against a separate group of users who are not served the ads. Because this is conducted at the user level, lift tests typically observe platform metrics. Google Ads’ Brand Lift and Conversion Lift studies observed in Video-based campaigns are very common examples.
H3: Get the Timing Right
External factors can influence an incrementality test, so it’s important to design a test in a way that mitigates these concerns. Things like sales or promotions, seasonality, anticipated shifts in competition, and even changes to internal operations can all sway performance in a given period. Sticking with the ecommerce example above, some of the things that might sway results could include things like shipping delays and products going out of stock. And it isn’t only the testing period that’s important.
Pro-Tip: A comprehensive incrementality test will include a built-in halo period after the test is concluded, in which overall performance is observed to see if there is a significant change within the test group after moving back to normal marketing activities.
In addition to calendar timing, one of the major considerations for an incrementality test is determining how long it should run. Obviously, budget will usually play a role in deciding the duration of a test period, but stakeholders must also agree on a minimum detectable effect (MDE). The MDE is the smallest amount of performance fluctuation needed to convince all parties that the results of an incrementality test are conclusive. In plainer terms, how much impact must actually be observed to determine that the test results can be trusted? And to avoid confusion, remember that MDE is not the same thing as statistical significance. MDE is estimated before a test is launched, but statistical significance is determined after it concludes and is based on the data collected.
MDE will also vary based on the sample size being observed, and the larger the affected population, the greater the confidence in the results of a test. This is why it’s important for advertisers to estimate the audience size and to agree on what constitutes an actual change in performance that is most likely a result of the test as opposed to simple period-to-period fluctuations in business.
Clarify Success Details
Reporting might be the most important element of any incrementality test, especially when it comes to getting buy-in and maintaining trust. When explaining the recommended nature and structure of a test, it’s not uncommon for stakeholders to voice objections right out of the gate. Anyone who’s pitched a geo holdout test has probably heard something like “Turning off ads in ten states sounds like losing money on purpose!”
That’s why advertisers must be very clear about articulating any anticipated risk, reiterating the MDE for achieving statistical confidence in the test, and providing an estimate of how long it might take for performance to normalize once a test is concluded. And as any skilled Paid Search manager knows, there must be a regular reporting cadence to ensure the test is being constantly monitored. These things are essential to getting approval from broader teams.
Interpreting and Reporting Results
Because the KPI and source of truth were agreed upon before the test was ever launched, there shouldn’t be any surprises when it’s time to review test results. To do this effectively, advertisers should restate the anticipated MDE and identify the actual change that was driven during the testing period.
For example, if the example ecommerce company above does choose to conduct a geo holdout test by pausing Non-Brand Search campaigns in ten states, then the results should explain what happened in that test group as well as the control group.
To tell that story effectively, results must include the following information:
- Overall KPI in the test locations during the testing period
- In this example, overall Shopify sales in those states might be the agreed-upon KPI.
- Overall KPI in non-test locations during the testing period
- Overall KPI in test locations in the pre-period leading up to the test
- Overall KPI in non-test locations for the pre-period leading up to the test
- Pre/post KPI delta in test locations
- Pre/post KPI delta in non-test locations
- Overall KPI in test locations after conclusion of test
Observing such a comprehensive series of data points allows the advertiser to show a well-rounded summary of the test’s impact. If the campaigns being paused legitimately drive business objectives, then pausing them in the test locations will lead to an observable decrease in Shopify sales during the test period compared to the pre-period. And if the drop is driven by pausing Non-Brand Search, the non-test locations are not likely to experience a similar drop.
The difference in performance between the two location sets is a strong indicator that Non-Brand Search is driving real incremental value to the bottom line, even if the campaigns don’t always receive full attribution for their contributions. If sales pick up in the test locations shortly after ending the test, this is additional evidence that Non-Brand Search is responsible for Shopify sales.
Inconclusive or Negative Results
There will likely be times when incrementality tests show that campaigns are ineffective, or at least not driving a significant enough result to move the needle. This is where strong advertisers have a chance to shine as great partners. When unfavorable results occur, it’s important to be honest about them, explain them clearly, and recommend an action plan.
What if pausing Non-Brand Search campaigns in the example above didn’t have a significant impact on performance? The data should show that in a concrete way, and an advertiser should feel comfortable exploring why. Maybe the campaigns focus on the wrong keyword themes, or maybe the advertiser is saturating the market so heavily through other tactics like PMax, Meta, and ChatGPT ads that Non-Brand Search doesn’t break through the noise. Whatever the case, exploring the “why” opens the door for additional testing and optimization.
Action plans are built off things that aren’t immediately successful, which is why it’s worth developing an ongoing testing framework as opposed to treating each test as a separate one-off.
Incrementality Testing as a Practice
Incrementality testing is not about finding a single clear answer based on one snapshot in time. It should be an ongoing behavior that constantly makes accounts more effective. There is no perfect system for measuring the effectiveness of any channel in today’s omnichannel environment. Still, a strong incrementality testing framework creates confidence, keeps strategies fresh and relevant, and above all else, it ties Paid Search efforts to greater business objectives. It can easily stand as the cornerstone of an effective account management framework. No one hires a Paid Search manager or agency because they can achieve stronger click-through rates, but rather because they anticipate that the partnership will help their business. Incrementality testing is the secret weapon that gives life to that assumption.
