Stache n' Scale

Metrics

August 17, 2026

Cut the Muslin First: Incrementality Testing That Ends the Attribution Argument

Attribution assigns credit after the fact. It never observes the world where your ads were absent, which is how platform-reported return keeps flattering channels that are not earning their keep. Incrementality testing withholds the treatment and reads the difference. Here is how to design holdouts and geo lift tests that survive scrutiny.

By Obert Kong

Growth Architect

A tailor cuts a muslin before touching the good wool. Do the same before you cut the budget.

Every growth team has a version of the same Monday argument. The platform dashboard says retargeting returned four dollars for every one. The finance model says the company grew about as fast in the months when retargeting was switched off. Two sincere numbers from the same weeks, and nobody can say which one deserves next quarter's budget.

A tailor does not settle a fit argument by staring at the pattern. The tailor cuts a muslin, a cheap test garment in plain cloth, pins it on the body, and finds where the shoulder pulls before touching the good wool. The muslin is disposable, honest, and built to answer one question at a time.

Incrementality testing is the muslin. It does not ask which touchpoint stood nearest the conversion. It asks the harder question: if we had not run this, what would have happened anyway? You answer by withholding the treatment from a comparable slice of the world and reading the difference. Everything else in measurement is inference. This is evidence.

What follows: the three designs you can actually run, a worked geo holdout, the failure modes, and a Monday cadence.

Attribution tells you who was standing nearby. Incrementality tells you who did the stitching.

— THE SCALE MANIFESTO, 1924 (REV. 2024)

Why Attribution Cannot Settle the Budget Question

Attribution is bookkeeping. It observes conversions that already happened and distributes credit across the touchpoints it can see, according to rules somebody chose. That is useful for optimizing inside a channel, because relative movement in a consistent system carries signal. It is useless for deciding whether the channel deserves to exist, because it never observes the world in which the ads were absent.

The research here is not subtle. In a comparison of measurement approaches across fifteen large field experiments at Facebook, Gordon, Zettelmeyer, Bhargava, and Chapsky ran roughly 500 million user-experiment observations and 1.6 billion impressions through both randomized experiments and the observational models advertisers actually use. The observational models frequently failed to recover the experimental effect, even after conditioning on extensive demographic and behavioral data. Sometimes they overstated badly, sometimes they understated. The unpredictable direction of the error is the part that should worry you.

The eBay paid search experiments run by Blake, Nosko, and Tadelis are blunter still. Brand keyword ads showed no measurable short-term benefit: people searching for eBay found eBay. On non-brand keywords, new and infrequent users did respond, but most spend landed on frequent users whose behavior did not change, dragging average returns negative. They learned this by switching bidding off across a randomly assigned share of United States traffic and watching for what failed to happen.

Read that correctly. The lesson is not that paid media fails. The lesson is that effect size is an empirical question, specific to your business, your audience mix, and your current spend level, answerable only by withholding. Size channels off platform-reported return and you are pricing your budget with a number the platform has every incentive to keep generous. Our earlier note on CAC, LTV, and ROAS lays out the arithmetic. Incrementality is what makes the arithmetic true.


Three Cuts: The Designs You Can Actually Run

There are only a few ways to withhold a treatment. Which one fits depends on what you can randomize.

The user-level holdout

Where you control delivery to identified people, suppress a random slice of the eligible audience. Lifecycle email is the easiest start: exclude ten to twenty percent of a segment, keep them excluded for a full conversion cycle, then compare. Platform-native studies work the same way underneath. Meta's lift study API has you declare cells with a treatment percentage and a control percentage that sum to one hundred, and the platform handles randomization and suppression.

Strength: clean randomization, tight intervals, low cost. Weakness: it only measures what the platform can see and suppress, so it cannot tell you whether the channel is harvesting demand your brand already created.

The geo holdout

When you cannot randomize people, randomize places. Google's geo experiments methodology assigns non-overlapping regions to treatment or control and fits a linear model to estimate incremental return on ad spend. Google Ads productizes this as conversion lift based on geography, which uses Google Marketing Areas as experimental units and runs a contamination model to reduce the bias created when somebody sees an ad in one region and converts in another. The documentation is refreshingly candid that this design needs a bigger budget than the user-level version to reach significance, and it makes you pass a feasibility check first.

If you have too few regions for region-level regression, the time-based regression approach predicts a counterfactual time series for the market instead, and the trimmed match design tightens precision on paired geo tests. GeoLift brings synthetic control methods to the same problem in R, with diagnostics that flag when your matched markets are not actually matched.

Strength: it works where no reliable user-level identity exists, and it captures total business effect including cannibalization. Weakness: expensive in foregone revenue, slow to read, and fragile if your regions were never comparable in the pre-period.

The on and off test

Switch the channel off for a fortnight and watch the line. These are the blunt shears. Nothing separates your intervention from seasonality, a competitor's launch, or a press cycle, so a single on and off read is a tripwire, not a verdict. Use it to decide whether a proper experiment is worth designing, never to defend a budget.


A Worked Muslin: Sizing a Geo Holdout on B2B Pipeline

Take a B2B software company at twenty-four million dollars in annual revenue. Paid media runs eighty thousand a month across search and social. The dashboards report a blended four point two return. The sales cycle averages forty-five days. Leadership wants to double spend next quarter, and nobody can say what the last dollar bought.

Design decisions, in order:

Now the arithmetic. Across the six-week window, treatment regions spent one hundred twenty-four thousand dollars more than control regions. Treatment regions produced one million four hundred ten thousand dollars in qualified pipeline. The counterfactual, predicted from the matched control regions and the pre-period relationship, was one million one hundred fifty thousand. Incremental pipeline: two hundred sixty thousand dollars.

That is an incremental pipeline return of 2.1 on spend, which sounds tolerable until you convert pipeline into money. At the historical win rate of twenty-six percent, two hundred sixty thousand in incremental pipeline is about sixty-seven thousand six hundred dollars in incremental revenue against one hundred twenty-four thousand in incremental spend. First-order incremental return: 0.55. The platform reported 4.2. The experiment said the marginal dollars were burning gross profit.

One more number matters more than all of the above. The ninety percent confidence interval on incremental pipeline ran from ninety-five thousand to four hundred twenty-five thousand dollars. That is a loose garment: it rules out the platform's story but does not support a precise reallocation. Because the decision rule was written before the data arrived, the team held to it: cap paid at current spend rather than cut to zero, then run a spend-level test next quarter to find where returns bend. A capped channel with a known range beats an uncapped channel with a fantasy.


Failure Modes That Quietly Invalidate the Result

Underpowered by design

The most common failure is a test that could never have detected the effect you cared about. Do the power math first: conversions per cell per week, the minimum effect you would act on, and the weeks that combination requires. If the honest answer is thirty weeks, you do not have an experiment, you have a wish. Practitioner guidance for platform lift studies converges on a floor of roughly fifty to one hundred conversions per cell per week over two to four weeks of stable delivery. Most B2B teams cannot clear that on closed revenue, so they move the outcome up the funnel, or the unit from users to regions.

Contamination across the seam

People travel, devices multiply, and companies keep offices in several cities. When treated exposure produces control-side conversions, measured lift shrinks toward zero and you conclude the channel does nothing. Google builds a contamination model into its region selection because this bias is structural rather than rare. For account-based programs, randomize at the account level, never the contact level, or your control group fills with colleagues of treated buyers.

Meddling and peeking

Every mid-test optimization is a knife through the muslin. A creative refresh, a bid strategy swap, a budget reallocation, a campaign restart that triggers a fresh learning phase: each breaks the comparison you paid for. Peeking is the subtler version. Check daily, stop when the number looks flattering, and you have manufactured a false positive out of process rather than data.

The wrong outcome and the wrong window

Measuring demo requests in a business that lives on closed revenue flatters every top of funnel channel you own. Measuring closed revenue inside a six-week window in a business with a four-month cycle flatters nothing at all, and gets a good channel killed. Match the outcome to the decision, match the window to the cycle, and state both in writing before you start.

Treating one read as a law

An incrementality estimate is perishable. Creative fatigues, auctions reprice, the competitive set shifts, and baseline demand changes as the brand grows. Stamp every result with a date and an expiry, then retest the channels carrying the most budget at least twice a year. A two-year-old lift number is not evidence, it is folklore.

An underpowered experiment is not evidence. It is an expensive shrug.

— THE SCALE MANIFESTO, 1924 (REV. 2024)

The Monday Operating Cadence

Experiments fail as a capability long before they fail as statistics. They die because nobody owns the calendar. Put the discipline on a weekly cadence and it survives a quarter.

One caveat for the current search environment. A growing share of influence now happens where no click exists to withhold, making clean holdouts on organic and AI answer surfaces close to impossible. Our note on measuring growth when nobody clicks covers the proxies that survive there. Run experiments where you can withhold, use disciplined proxies where you cannot, and never let the two share a column as if they carried equal weight.


Tooling Without the Theater

Tooling is no longer the constraint: platform lift studies, GeoLift, and matched markets are all within reach of a small team. The safeguard that matters is procedural, not technical. Write the hypothesis, the power math, and the decision rule in one document before the test starts, and have the budget owner sign it. A pre-registered decision rule is the only thing that stops a wide confidence interval from being read as whatever the loudest person in the room already believed.


The Checklist Before You Cut

Cutting a muslin costs cheap cloth and a week of patience. Cutting the good wool from a pattern nobody checked costs the season. Attribution is the pattern: worth having, never worth trusting on its own. The withheld cell is the only place your growth model gets fitted to the actual body of the business.

#Incrementality Testing#Geo Experiments#Marketing Measurement#Holdout Tests#Paid Media#Attribution#iROAS