You turn off a $9,000-a-month Meta campaign for two weeks as a test. Revenue barely moves. You turn it back on, relieved, and tell yourself the test was inconclusive because two weeks is not long enough.
It was long enough. It just answered a question you were not ready to hear: a meaningful share of that spend was not creating sales, it was claiming credit for sales that were going to happen anyway.
That is the gap between what your ad platforms report and what incrementality testing measures. Platform ROAS tells you what got credited. Incrementality testing tells you what actually changed because the ad ran.
- Incrementality testing measures the sales an ad campaign actually caused, not the sales it got credited for, by comparing a group exposed to the ad against a comparable group that was not.
- Meta's own eligibility bar for a Conversion Lift study, roughly $5,000 in spend and 500 conversions in the test window, is out of reach for a lot of single campaigns, which is why the native tools quietly go unused by the brands that need the answer most.
- A manual budget on/off test measured against blended revenue, not platform-reported conversions, is a realistic substitute at almost any spend level, if you run it for long enough and read the result correctly.
What incrementality testing actually measures
Every ad platform will hand you a conversion number and a ROAS. Neither one tells you what would have happened if the ad had never run, and that missing comparison is the entire point of incrementality testing.
Google's own measurement guidance frames modern measurement as three separate tools that answer three separate questions: attribution explains which touchpoints get credit for a conversion that already happened, marketing mix modeling estimates the aggregate effect of each channel over time, and incrementality testing isolates cause and effect directly by holding part of the audience back from the ad. None of the three replaces the other two, and a store relying only on attribution is answering "who gets the credit" while never asking "did this spend cause anything."
Measured's comparison of attribution and incrementality puts it plainly: attribution is a bookkeeping exercise applied after the fact, while incrementality is a controlled test that isolates the causal effect of the spend itself. A campaign can have a clean, defensible attribution model and still be spending against demand that was never actually created by the ad.
Why platform ROAS overstates what your ads are doing
In a delivered paid media audit, Meta reported a 4.2x ROAS and Google reported 3.8x for the same account in the same window. Blended revenue against total ad spend, the number that reconciles against what the business actually took in, worked out to roughly 1.9x. Both platforms were technically reporting correctly by their own rules. Both were also counting the same customer's purchase more than once, and neither number reflected how much of that revenue would have shown up with the ads turned off.
This is not a tracking bug to fix. It is what happens when every platform is incentivized to claim credit for as much of the funnel as its attribution window allows.
Meta's Conversion Lift methodology works by randomly splitting the audience into a group that sees the ad and a holdout group that does not, then comparing actual purchase behavior between the two. Meta's own eligibility requirements ask for a minimum of $5,000 in spend and roughly 500 conversions with a 1-day click, 7-day click, or 1-day view attribution setting before it will call a result statistically reliable. Google's geo-based Conversion Lift studies do the same comparison across geographic markets instead of individual users, and report the result as incremental return on ad spend, distinct from the platform-attributed ROAS on the same campaign.
The volume problem nobody mentions upfront
Here is the part that gets skipped in most explanations of incrementality testing: Meta and Google's native lift tools are built for accounts with enough conversion volume to detect a statistically stable effect, and a large share of ecommerce advertisers never reach that bar on a single campaign.
Meta's own algorithm guidance recommends a minimum of 25 to 50 conversions per week per campaign just for stable delivery, well below the roughly 500 conversions Meta asks for before it will call a Conversion Lift result statistically reliable. A store running one campaign at a $40 cost per purchase would need to spend around $20,000 in the test window just to generate that many conversions, on top of clearing Meta's own $5,000 spend minimum. That store is not doing anything wrong. It is simply too small for the tool that everyone assumes is the default way to measure incrementality.
This is the gap a lot of incrementality-testing content quietly skips past, because most of it is either written for enterprise brands with the budget to run a formal geo-holdout, or it is a platform's own case study assuming you already qualify. A brand well under that spend level needs a different starting point.
What to run instead, at almost any spend level
If your account is not large enough for a native lift study or a formal geo-holdout, a manual on/off test is still a real incrementality test, as long as you measure it correctly.
- Pick one channel and one metric. Choose the campaign or platform you most suspect is overstating its impact, and commit to blended store revenue (from Shopify or your order system, not the platform dashboard) as the read, decided before the test starts.
- Turn it fully off, not down. Reducing budget by half muddies the comparison. A clean off period removes the variable entirely.
- Run it for at least three to four weeks. Meta's own setup guidance recommends a minimum 28-day duration to reach statistical significance, and Google's lift-testing guidance points to a similar multi-week window. A shorter test invites exactly the "inconclusive" reading that sends people back to the platform dashboard for reassurance.
- Compare against a same-length prior period and, where possible, a control segment or geography that kept running. If you have distinct enough regions or customer segments, holding one back while the rest continues gives you a cleaner comparison than a single before-and-after read.
- Decide your sample size and duration before you start, not after you see the number. The same discipline that prevents an underpowered A/B test applies here. Our breakdown of A/B test sample size for ecommerce covers the math for deciding how long a test needs to run before you can trust the result, and the same logic applies to a spend holdout.
None of this requires a data science team or a six-figure incrementality platform. It requires turning something off, watching the number that actually reflects revenue, and resisting the urge to call it after five days.
Reading a null or negative result without panicking
A null result means blended revenue did not move when the spend stopped. A negative result means the holdout group did roughly as well or better than the exposed group. Neither means the channel is worthless.
Both usually mean one of two things: the ad was reaching people who were already going to buy (a brand-search or retargeting campaign capturing existing demand instead of creating it), or the budget level tested was past the point where more spend produces more sales. This mirrors a pattern we cover in our piece on ecommerce attribution models: a channel that looks strong under last-click credit can still show weak or negative incrementality, because attribution and causation are measuring two different things entirely, and a channel can score well on one while failing the other.
The fix is rarely "turn it off forever." It is usually "change the targeting, the audience, or the budget tier, then retest," the same way you would iterate on a losing landing page rather than declare the entire funnel broken from one result.
What to actually do with this in the next 30 days
- Pull blended revenue and spend by channel for the last 90 days, and compare it against what each platform reports for the same window. A gap of a full turn of ROAS or more, the kind we found in the 4.2x-versus-1.9x example above, is your signal that a test is overdue.
- Pick the single campaign most likely to be claiming credit it did not earn, usually branded search, retargeting, or a Performance Max campaign with a high brand share.
- Turn it off completely for three to four weeks and track blended revenue daily against the same period a month earlier.
- Read the result against the decision you made before starting, not against whatever number feels most comfortable afterward.
- If the volume and budget support it, layer in Meta's or Google's native lift tool for a second, platform-native confirmation once you already have a manual baseline to check it against.
If your reconciliation step turns up numbers that will not tie back to store revenue at all, that is usually a sign the measurement needs attention before a test result can be trusted. A paid media audit is built to find exactly that kind of discrepancy first, so the incrementality test you run afterward is measuring a real effect instead of chasing a number that was already broken. For stores managing attribution across more than one platform, our marketing analytics services build the blended reporting layer this kind of test depends on, so blended revenue is a number you can pull in minutes instead of reconciling by hand every time you want to check one.
Platform ROAS will keep telling you the ads are working. Incrementality testing is the only way to find out if that is true.

