On A/B Testing
The stronger a platform's machine learning gets, the more A/B testing comes down to the brand knowledge behind deciding what to compare.

Launch a campaign today and the platform finds the audience, sets the bids, and shifts budget toward whichever creative performs — all on its own. With fewer levers left to pull by hand, we keep hearing the same question: doesn't the platform run the A/B test for us now? The answer is the opposite. The stronger the machine learning gets, the more results depend on the person deciding what to compare — the creative team and marketers who understand the brand.
What changes about A/B testing as machine learning gets stronger?
The variable you can test shifts from "who" to "what." With campaigns like Meta's Advantage+ Shopping and Google's Performance Max — which set targeting, bidding, and placement automatically — now the default, splitting audiences to compare them leaves little room. The largest input a marketer still controls is the creative itself.
That does not mean uploading more assets is enough. The platform can only choose from the candidates a marketer supplies. The quality of that pool is the ceiling on performance, and designing the pool is still a human job.
| Category | Manual media buying | Machine-learning automation |
|---|---|---|
| Main test variable | Audience, bid, placement | Message, value proposition, creative format |
| Who optimizes | The marketer, by hand | The platform algorithm, in real time |
| The marketer's core role | Setup and monitoring | Hypothesis design and building the candidate pool |
| Bottleneck on results | Headcount and time | Ideas worth comparing |
Why is the "winning" creative the platform picks not always the right answer?
Because machine learning optimizes toward the signal you train it on — usually a short-term metric like a click or a purchase. The tone a brand has built over years, and how people think about it long term, are not in that signal. So provocative copy and heavy discount messaging tend to win on short-term conversion, and following those results indefinitely leaves you with a brand people respond to only when it discounts.
Comparing creatives inside one ad set is not an experiment
Put several creatives in a single ad set and the platform quickly concentrates budget on whichever gets early traction. A creative that barely received impressions did not lose; it was never evaluated. To judge creatives against each other, you need the split-testing feature that divides budget and audience into controlled cells.
Machine learning picks the best of the options it is given. It does not create the options.
So what should you actually be testing?
Not button color, but what the brand promises its customers. Minor variables get sorted out by the algorithm, but which promise sounds like this brand is something only people who know the brand can decide. A good test is one whose result changes the direction of the next round of creative.
| Category | Weak hypothesis | Strong hypothesis |
|---|---|---|
| What is compared | Red button vs. blue button | "Price" message vs. "saves you time" message |
| Starting point | Whatever is easy to change | Why customers choose us |
| Use of the result | Swapping one asset | Setting next quarter's message direction |
| Brand criteria check | None | Applied as a filter before testing |
Brand understanding matters at two points: coming up with alternatives worth comparing, and ruling out creative you should not run even when it wins on the numbers. Neither judgment exists in platform data.
How do you turn brand understanding into test design?
Start by writing down your brand criteria before any test goes live. Read the results without them and you will follow whichever number looks better; do that long enough and the brand starts to resemble the platform's algorithm.
- 01
Define the brand criteria
the tone to protect, the language you will not use, and the core promise to customers, on a single page.
- 02
Design the hypothesis
pick two or three competing value propositions drawn from why customers choose you.
- 03
Run it under control
use split testing to separate budget, duration, and audience.
- 04
Interpret and apply
scale only when you can explain the win in the brand's own language.
That last step matters most. If you cannot say why the winning creative won, the result leads nowhere. Only an explanation turns a single win into knowledge the team keeps.
What was different in an actual campaign?
In a campaign for a large e-commerce brand, we spent two weeks comparing eight value propositions, and conversion rose by 28%p. The budget and the flight were unchanged; the only difference was what we put up for comparison.
Those eight were not the ones that looked most likely to perform — they were narrowed down from the brand's own criteria.
28%p
change in conversion rate
2
weeks of testing
8
value propositions compared
If you want to design creative tests that start from brand criteria
get in touch with CREATIP- A/B testing
- creative testing
- performance marketing
- ad optimization
