Never Always, Never NeverNever Always,
Never Never
The BookAI CreativeAI ProjectsChatNewsletter
Buy the Book
Never Always,
Never Never.

Strategic Marketing in an AI World.
By Patrick Gilbert.

Explore

  • The Book
  • The Author
  • Free Chapter
  • Buy
  • Resources

Connect

  • Newsletter
  • Learn
  • AI Projects
  • Blog
  • Chat with the Book
  • Subscribe

Stay Connected

Get updates, bonus frameworks, and new AI project showcases.

© 2026 Patrick Gilbert. All rights reserved.

AdVenture MediaContact
Measurement7 min readJuly 31, 2026

Incrementality Testing: Your Platform ROAS Is Lying to You

Patrick Gilbert

Patrick Gilbert

CEO of AdVenture Media. Author of Never Always, Never Never.

Platform-reported ROAS overstates true incremental impact by 20% to 60%. That gap isn't a rounding error. It's the difference between marketing that causes growth and marketing that merely correlates with it.

According to EMARKETER and TransUnion data from July 2025, 52% of U.S. brand and agency marketers now use incrementality testing to measure campaigns. Another 36.2% plan to invest more in it over the next 12 months. The reason is simple: the industry has finally acknowledged that attribution dashboards tell you what happened, not what your ads caused. Those are not the same thing.

Patrick Gilbert makes this argument in Never Always, Never Never, specifically in the measurement chapters: attribution is an execution tool, not a verdict. Incrementality testing is the closest marketing gets to a true experiment, and treating anything less as definitive is how budgets get misallocated for years.

What Incrementality Testing Actually Measures

Most measurement tools track what happened. Incrementality testing asks a different question: what would have happened without the ad?

Straightforward in method, the approach holds out a control group that doesn't see the marketing, exposes a treatment group that does, and measures the difference in conversions, revenue, or whatever outcome you care about. That difference is the incremental lift, the behavior your campaign actually caused, rather than behavior that would have occurred anyway.

Counting on attributed conversions to reflect true impact is risky, because a meaningful share would have happened organically. Meta lift studies show that 10% to 40% of attributed conversions fall into this category. Retargeting is a particularly common failure point: according to industry summaries, 20% to 30% of retargeting conversions would have happened without the ad. These people were already going to buy. You paid to interrupt them on their way to the checkout.

The difference between incrementality and attribution is not academic. It is where budget allocation decisions either get smarter or stay broken.

Why Attribution Gets This So Wrong

Attribution assigns credit across the digital touchpoints that preceded a conversion: clicks, impressions, emails, retargeting, paid search. It is useful for tactical decisions within a channel. It is dangerous as a foundation for strategy.

Structurally, attribution systematically overvalues the channels that are easiest to track, especially bottom-funnel tactics. Google Search and retargeting look efficient in every attribution model because they touch users at the exact moment of intent. But a user searching for your brand name was often already going to buy. Stella's benchmark data makes this concrete: across 225 incrementality tests, branded Google Search produced a median iROAS of 0.70x. That means for every dollar spent on branded search in that dataset, the true incremental return was less than a dollar. The channel looked efficient in attribution. The experiment said otherwise.

In Never Always, Never Never, Gilbert uses the analogy of an NFL offensive lineman. Attribution rewards whoever scores the touchdown. It struggles to value the players who created the conditions for the score, the blockers, the route runners, the quarterback who extended the play. Upper-funnel brand campaigns work the same way. They rarely appear in attribution models as the final touch. But Stella's benchmark median iROAS of 2.31x across all tested campaigns suggests that, when you actually measure causal lift, many of those less-visible investments are earning their keep.

Les Binet and Peter Field made this core critique in Marketing in the Era of Accountability, their analysis of more than 1,000 campaigns from the IPA Effectiveness Awards DataBank. Narrower measurement leads organizations to optimize for the wrong things: short-term, attributable, visible gains at the expense of compounding long-term growth. ROAS as a primary optimization target is a particularly reliable path to this outcome.

The Three-Tool Framework (And Where Incrementality Fits)

Testing for incrementality does not replace attribution or marketing mix modeling. It fills a specific gap that neither tool can fill on its own.

Marketing mix modeling vs attribution is a useful place to start. MMM operates at altitude. It looks across your full channel mix over time and asks which combinations of investment tend to correlate with business outcomes. It is a budgeting and planning tool, useful for deciding whether paid social should represent 20% or 40% of total spend, not for deciding which creative to pause this week. Attribution works at the opposite level: individual touchpoints, near real-time signals, tactical optimization within a channel.

Incrementality sits between them as the validation layer. It pressure-tests the assumptions that both MMM and attribution produce. When your MMM model suggests branded search is a strong performer, incrementality testing can tell you whether that relationship is causal or correlational. When your attribution dashboard shows retargeting generating strong ROAS, a holdout test can tell you how much of that would have converted anyway.

Used in sequence, the three tools form a measurement system. MMM sets direction and allocation guardrails. Incrementality validates the causal claims. Attribution provides fast signals for in-channel optimization. Each one makes the others more trustworthy when they are used for the right questions.

Gilbert covers this framework in the book under the lens of what he calls scoreboards versus film rooms: attribution, MMM, and incrementality are film room tools, built for learning and optimization. Problems start when organizations treat them as scoreboards, final verdicts on whether a campaign, channel, or agency is "good." When survival depends on producing attributable ROAS, the incentives drift toward whatever is easiest to measure and easiest to claim credit for. That is a framing problem, not a data problem.

Retail Media Is Forcing the Issue

Among the sharpest near-term pressures on incrementality adoption is retail media. According to the ANA, 71% of advertisers now rank incrementality as their single most important KPI in retail media. The reason is structural: retail media networks sell placements based on closed-loop attribution against their own sales data. That attribution looks clean. It is not necessarily causal.

Buyers need to know whether their sponsored product placements are driving purchases that would not have happened otherwise, or whether they are paying for visibility against shoppers who were already going to buy that product in that session. An attribution model built on the retailer's own data cannot answer that question neutrally. An incrementality test can.

Retail media has become a catalyst for broader incrementality adoption as a result. Measurement stakes are high, the attribution conflict of interest is visible, and the question of causal lift is concrete enough that finance teams understand it.

The Cost Barrier Is Shrinking

For years, incrementality testing was a tool for large advertisers. Minimum investment required to run a properly powered geo lift test or holdout experiment was substantial. One vendor-cited figure puts the historical floor at roughly $100,000 per test, with the current floor closer to $5,000 following Google's expansion of lower-budget testing access. That figure comes from vendor commentary and has not been independently audited, so it should be treated as directional rather than definitive. But the directional claim is credible: the cost of running experiments has fallen, and access has expanded.

Mid-market brands that previously had no practical way to validate whether their campaigns were driving incremental behavior can now, in principle, run a test that answers the question most worth answering: is any of this causing growth? A brand spending meaningfully on paid media is no longer necessarily priced out of that answer.

Measurement vendors including Measured, Skai, and Stella have also built tooling that simplifies test design and analysis, which lowers the operational barrier alongside the financial one. Industry summary data cited by EMARKETER suggests this is translating into adoption: 27.6% of U.S. brand and agency marketers say expanding incrementality testing is a top measurement priority.

At AdVenture Media, the practical push toward incrementality has mirrored what the industry data shows: the question is no longer whether to test causality, but how to build it into a regular measurement cadence rather than treating it as a one-off project.

The Limits of the Method

Incrementality testing is the closest marketing gets to a controlled experiment. It is not a controlled experiment.

Results are influenced by timing, creative, competitive activity, and market conditions at the moment the test runs. A single holdout test produces one data point under one set of conditions. In academic research, meaningful conclusions require replication across multiple runs to understand variance and statistical significance. Marketing teams rarely do this. They run a single test, get a number, and treat it as definitive truth, which is exactly the same mistake they were making with attribution.

Geography-based tests introduce additional complexity. If your treatment and control regions have different baseline conversion rates, different competitive dynamics, or different seasonality patterns, the result will reflect those differences as much as the actual campaign effect. Good test design requires pre-test validation, matched market selection, and honest assessment of whether the results are statistically significant or just noise.

Honestly summarized, Gilbert quotes statistician George Box in Never Always, Never Never: all models are wrong, but some are useful. Incrementality testing is useful. It is more causally rigorous than attribution. It is not a final answer. The right question is never whether the model is true. It is whether the model is useful for the decision you are trying to make.

For budget allocation decisions and channel validation, incrementality is the most useful tool available. For in-week creative optimization, attribution is more practical. For annual budget planning, MMM provides the broader context neither tool can offer on its own. The measurement system is the combination, not any single method.

What This Changes in Practice

If your current measurement stack is built primarily on platform attribution, the gap between what you think your campaigns are returning and what they are actually causing is probably material. Industry summaries cite that gap at 20% to 60% against platform-reported ROAS. That range is wide, but even the low end represents a meaningful misallocation of budget.

A practical starting point is identifying which channels and tactics have the highest plausible gap between attributed performance and incremental performance. Branded search is a strong candidate, as Stella's benchmark data suggests. Retargeting is another, given the organic conversion rates in that audience segment. These are channels where attribution tends to claim credit for intent that already existed.

Running a holdout test on one of those channels does not require a large budget or a dedicated analytics team. It requires a clear hypothesis, a properly matched control group, and enough patience to let the test run long enough to reach statistical significance. The how to measure marketing effectiveness framework covers the right sequencing.

Beyond methodology, the broader shift is mental. Incrementality testing forces you to treat measurement as a learning system rather than a scoreboard. You are not running a test to prove a channel works. You are running a test to find out whether it works, under what conditions, and by how much, and then using that information to allocate the next dollar more honestly than you allocated the last one.

That is the difference between accountability theater and actual accountability. The dashboards look the same. The decisions are very different.

Patrick GilbertPatrick Gilbert

Patrick Gilbert is the CEO of AdVenture Media and author of Never Always, Never Never and the bestselling Join or Die. He has been ranked among the top 5 PPC experts worldwide and has delivered keynotes at Google events across three continents.

More about Patrick →

Enjoyed this?

Subscribe for more articles on strategy, AI, and what's actually working in marketing.

No spam. Unsubscribe anytime.

Keep reading

Measurement

Marketing Attribution Is Broken. Here's What to Do Instead.

75% of marketers say attribution underperforms. Stop using measurement tools as scoreboards and start treating them like film rooms.

Measurement

ROAS Is a Rearview Mirror: The Metric That Lies to Ecommerce Brands

ROAS measures credit assignment, not causal impact. Here's why the metric misleads ecommerce brands and what to use instead.

Strategy

Brand vs Performance: The Budget Split Question Has a Wrong Answer

The 60/40 rule isn't dead, but chasing a fixed split is. Here's how to actually allocate your budget between brand and performance marketing.