The Meta Ads Testing Rules We Use (And Why We Never Switch Ads Off)

Creative testing is where most Meta budgets get wasted, and almost never because the creative was bad. It gets wasted because the person running the account keeps intervening.
We run testing on two rules. Both are simple. Both are the opposite of what most founders' instincts tell them to do.
Table of Contents
- Rule 1 — the ±10% budget rule
- Rule 2 — ads are never switched off
- Some die, some come back
- How the testing stage is built
- What we do not do
Key Takeaways
| Rule | The action | What it prevents |
|---|---|---|
| ±10% on CPA | Beats target CPA: +10%. Misses target CPA: −10%. | Emotional budget decisions and delivery shocks |
| Never switch off | Let Meta allocate within the ad set | Killing ads on samples too small to judge |
| Don't chase revivals | An ad that comes back is normal, leave it | Resetting the learning that was working |
Rule 1 — the ±10% budget rule
Beats target CPA → +10%. Misses target CPA → −10%.
That is the entire rule, applied at the ad set level.
Three things make it work, and none of them are about the number ten specifically.
It is small enough not to shock delivery. Large budget swings push an ad set back into learning. A 10% move is usually absorbed without resetting anything, so you can apply it repeatedly without paying the learning tax each time.
It compounds. An ad set that keeps beating CPA gets more budget every day, and it gets it gradually enough that performance has a chance to hold as it grows. An ad set that keeps missing gets quietly starved rather than dramatically killed. Over a couple of weeks the budget has migrated to the right places without anyone making a big decision.
It ends the argument. This is the real value and the reason to write the rule down before you need it. Without a rule, every morning becomes a negotiation with yourself about whether the underperformer deserves one more day. With the rule, the number decides and you apply it. The discipline is not in having a clever threshold — it is in not overriding the threshold you set.
The obvious question: what is target CPA? It is whatever your margin says it can be, and it is the one number you should have decided before spending anything. If you do not have one, the ±10% rule has nothing to hold on to and you will end up back in the daily negotiation.
Rule 2 — ads are never switched off
Inside the testing ad set, we do not turn ads off. We let Meta sort them.
This is the rule people argue with, so here is the reasoning.
When you kill an ad on day two because it has not converted, you are making a statistical claim — that this creative will not hit target CPA — on a sample that cannot support it. A handful of clicks and no purchases is not evidence of a bad creative. It is the expected result of a small sample on any creative, including the ones that go on to carry your account.
Meanwhile Meta is already doing the job you are trying to do, with much better information. Within an ad set it allocates impressions towards whatever it estimates will perform, continuously. A creative that genuinely is not working will quietly stop getting delivery. You do not have to do anything for that to happen — and when you do intervene, you are overriding a system that sees the auction and you do not.
The practical version: put 2 to 4 creatives in an ad set, set the budget, and let allocation happen. Judge the ad set on CPA, not the individual ads on your feelings about them.
Some die, some come back
An ad flatlines for a week. You forget about it. Then it starts spending again and performing.
This is normal. It is not a bug, it is not a glitch, and it does not mean you have discovered something exploitable. Audience availability shifts, competition in the auction changes, seasonality moves — and an ad that was uncompetitive last Tuesday can be competitive this Tuesday without anything about it changing.
We don't chase them. Chasing means bumping the budget to capitalise, or editing the copy to "help it along," or duplicating it into a new ad set to give it room. All three do the same thing: they disturb the exact conditions that produced the recovery. The ad was working because of where it sat and what it had learned. Change either and you are starting over.
Leave it. If it holds, the ±10% rule will find it.
How the testing stage is built
For completeness, the shape the two rules sit inside:
- ABO, so each angle gets a controlled budget rather than competing for one pot
- 2 to 4 creatives per ad set
- Small daily budget per ad set
- Testing lives away from scaling — in its own campaign, or in its own ad set with spend limits if budget is too thin for two campaigns
Why the separation matters is the subject of the three-stage structure, and the specific layout for your budget is in structure by spend level.
The bit worth repeating here: if your test creative shares a campaign with a proven winner and there are no spend limits, none of these rules will save you. The winner takes the budget, your test never gets delivery, and the CPA you are reading is noise.
What we do not do
We don't judge on a good day. One strong Tuesday is not a promotion. The next stage exists precisely to check whether performance holds under real competition.
We don't edit ads that are working. Editing can reset learning. If you want to test a change, duplicate it and test the change — do not overwrite the thing that is currently paying for the account.
We don't rebuild the account when a week goes badly. Structure churn means nothing ever accumulates enough history to be judged. A bad week inside a stable structure is data. A bad week followed by a rebuild is nothing at all.
We don't test angles we cannot service. A winning ad pointing at a page that does not deliver the promise will underperform in ways that look like a creative problem. Meta's Estimated Action Rate treats the URL destination as one of its three inputs — the page is part of the ad, whether or not you treat it that way.
Applying ±10% by hand every morning across a dozen ad sets is exactly the kind of job worth handing to a machine. Here is how we run it with Claude, prompts included — or book a strategy call and we'll audit your testing setup on the call.