Craft
Targeting and bidding got automated. The operator job left is a muslin process for hooks and offers: isolate one variable in a dedicated test campaign so Advantage+ and PMax cannot starve it. The unit is the ad, not the channel holdout.
The creative Slack lands at 4:47 p.m. Forty exports in a folder labeled TEST_WAVE_12. No hypothesis. No control. No kill rule. Brand wants more UGC. Media wants more volume. Finance wants CPA to hold. By Friday the folder is empty, spend is up, and nobody can name one reusable rule the account learned.
A tailor does not pin forty scraps on a mannequin and call it a fitting. The sample rack holds one cut at a time, in cheap cloth, until the shoulder either sits or gets discarded. Creative testing is that rack. Scaling campaigns are the finished suit. Hang scraps on the suit and the machine will dress the history it already trusts while the new cloth starves.
This is not a channel holdout. Our earlier note on cutting the muslin first asks whether a channel deserves to exist. This note asks whether a new ad ever got a fair read. Targeting and bidding got automated. The operator job left is the unit: isolate one variable, fund the read, then kill or graduate on a written rule.
What follows is the operating system: why Advantage+ and PMax are not labs, the six cuts on the rack, a worked fiction for sizing budget, the failure modes that fake learning, and a Monday cadence that keeps the rack honest.
Do not hang scraps on the finished suit and call the delivery report a test.
— THE SCALE MANIFESTO, 1924 (REV. 2024)
Advantage+ and Performance Max optimize for outcomes they already understand. Drop a new concept into a shared automated budget next to a proven winner and the system routes spend toward history. Rotation is not isolation. A concept that never clears a spend floor did not fail. It never got fitted.
Meta's Andromeda retrieval layer makes the creative the first gate. The machine reads the ad to decide who is even eligible for ranking. An ad that never makes the shortlist turns bid strategy into theater. Google's PMax asset groups behave the same way under a different name: the system finds what already works and feeds it.
Landing-page CRO is a different fitting. Our earlier note on micro-adjustments to landing pages covers the page after the click. The sample rack is the ad that earns the click in the first place.
Build a dedicated test campaign and keep it off the scaling campaign. On Meta, give each concept its own ad set budget (ABO). Identical broad targeting, placements, and bid strategy. Equal daily budget. Broad is intentional: you are testing the cloth, not a lookalike you already believe. On Google, keep tests off PMax: one concept per asset group in a separate campaign.
The scaling campaign is Advantage+, CBO, or PMax. It receives graduates. It does not host the lab.
A concept is an angle or an offer: problem-first versus feature-first, guarantee versus discount. A variant is a rendering of that concept: a different first three seconds, thumbnail, or closing ask on the same claim.
Test concepts first. Then variants. Change one variable: hook, or offer, or format, or proof. A winner that changed the opening, the deal, and the aspect ratio taught you that a bundle worked once. You cannot reuse any piece of it.
Write delivery, spend, and conversion gates before anything goes live. Read them in order.
Delivery. Enough impressions to read the opening. Hook rate (3-second views ÷ impressions) diagnoses the first seconds. CTR diagnoses whether the message earned a click after attention. Published Meta ranges often treat 25 to 30 percent hook rate as a workable baseline. Treat bands as directional. Your own history is the control.
Spend. Do not judge CPA until the concept has spent on the order of two times target CPA and has cleared weekdays plus a weekend. Underfed tests stay inconclusive.
Conversion. Enough outcomes to talk about efficiency. A practical floor is roughly 50 conversions per concept for a CPA verdict, with a few days to clear learning. Below that floor you are reading delivery, not efficiency.
Three verdicts, written before launch: kill, recut, graduate. Flat does not ship.
Read in funnel order. Each metric names a different seam.
Our earlier note on CAC, LTV, and ROAS is the arithmetic the scorecard has to serve. A cheaper CPA on low-LTV buyers is not a win.
Fatigue is not a weekly law. It is fast enough that a monthly review is already late. Frequency climbs, hook rate softens, and last week's winner starts rationing delivery to people who already saw it. Two to four new concepts per week is a practical target for many brands, sized to the budget formula below. Starved tests are not velocity. They are inconclusive spend with a production calendar.
Weekly test budget = target CPA × conversions required for a read × concepts in the window
The share of total media is that number divided by this week's spend. Sometimes 8 percent. Sometimes 35. Cut concept count until the number fits. A slide that says "reserve 10 to 20 percent" is a sanity check after the formula, not an input.
The unit is the ad, not the channel holdout.
— THE SCALE MANIFESTO, 1924 (REV. 2024)
The numbers below are a worked fiction, not a client case.
A direct-to-consumer brand spends $80,000 a month on Meta. Blended CPA is $40. They want four new concepts a week and a CPA read, so they set the conversion gate at 50 purchases per concept.
Four concepts × 50 conversions × $40 = $8,000 a week. That is $32,000 a month, 40 percent of spend. The formula will not fund four CPA verdicts without starving scale.
They cut the queue to two concepts. $16,000 a month. 20 percent. Each test can hit the gate.
They run those two in a dedicated test campaign, each ad set on its own budget, broad targeting, one variable (the hook), identical offer and format. Decision rule written before launch: after the spend and conversion gates, keep any concept within 20 percent of target CPA and recut it. Kill the rest. Flat does not graduate. Winners move, as the same ad, into the Advantage+ scaling campaign. That campaign is not touched during the test window except to receive graduates.
That is a sample rack. The finished suits hang somewhere else.
New concepts sit under one shared campaign budget next to proven winners. The algorithm concentrates spend on history. The new tests never clear the spend gate. You conclude they failed. They were never fitted. Fix: each test ad set on its own budget for the lab.
You test 9:16 against 1:1 against 4:5 before you have a concept that can survive any frame. Format is a real variable. It is a later variable. Order: concept, then hook, then variation.
A slide said reserve 10 to 20 percent for tests. Ten percent of a $15,000 account cannot buy four CPA-significant tests. Size from CPA × conversions × concepts. Then look at the percentage.
You check daily, stop when the number looks flattering, and scale a day-two leader. Early leads reverse. The sibling sin is mid-flight meddling: a bid change, a budget swing, a creative swap that tears the comparison you paid for. Write the window. Read once at the gates.
Twelve renders of the same claim get labeled velocity. The machine learns one angle twelve times. Cap variants until a concept clears the conversion gate. Then recut the winner.
A concept clears the gates, then gets dumped into Advantage+ as a folder with no naming, no recut plan, and no kill rule once frequency climbs. Graduation is promotion plus ownership. If the winner has no owner after it leaves the rack, you lost the receipt.
Tests die as a habit before they die as statistics.
Incrementality still sits above this system. A sample rack that funds a channel with no incremental lift is a well-run lab on the wrong floor.
Cutting muslin costs cheap cloth and a week of patience. Hanging scraps on the finished suit costs the season. The platforms will keep automating the rest of the stack. Let them. Your job is the rack: isolate the variable, fund the read, kill what fails the fitting, and only then sew it into the garment you actually sell.
The decision this week is not whether creative matters. It is whether the next folder in Slack gets a hypothesis and its own rack, or whether you will spend another quarter reading an Advantage+ delivery report and calling it learning.