When A/B testing is the wrong tool

Steve New

We once had two possible names for a course and no particularly good reason to trust our own preference between them.

There was a related free event coming up, so we used its promotion to get some evidence. We split the first promotional email between the two event names while sending both groups to the same registration page. One version got more registrations from that email, and we used that result when choosing the eventual course name.

It was a useful test because the uncertainty was narrow and the result could inform a real decision quite quickly. It was also important not to claim more from it than we had learned. The experiment did not prove that one course name would produce more sales. It showed that, in that audience and promotional context, one framing attracted more registrations.

That was better evidence than debating which name sounded nicer.

A/B testing is very good at answering questions with that sort of shape. Plenty of conversion problems do not have it.

What decision will the test change?

Suppose a business thinks its sales page is underperforming and wants to test another headline.

That may be useful. Before spending traffic on it, though, I would want to know what uncertainty we are trying to resolve.

If there is some evidence that the headline is causing confusion or failing to communicate something important, comparing alternatives makes sense. If serious prospects already understand the offer but keep asking whether the programme suits their experience or whether they can realistically attend, the headline may not be the most useful uncertainty to investigate.

A technically sound experiment can still spend weeks answering a fairly unimportant question.

I find it useful to ask what happens after each possible result. If A wins, what would we do? If B wins, what would we do? What if there is no meaningful difference?

If the action would barely change, the experiment probably does not deserve much traffic or attention.

Low-volume businesses have to choose their tests carefully

Large ecommerce businesses can investigate fairly small differences because thousands of relevant customers may pass through the same step.

A professional training programme can be in a very different position. It may have only a few hundred serious prospects for an intake, while a consultancy or expensive B2B service may have fewer genuinely comparable opportunities again.

If it would take several months to learn whether a minor wording change improves conversion slightly, the real question is whether that uncertainty deserves several months of customer behaviour.

Sometimes it does. Often there is a larger uncertainty nearby.

Recent buyers and serious non-buyers may be telling you the same thing. Support questions may repeatedly cluster around one point. Transaction or behavioural data may show that people progress normally until a particular stage.

None of that gives you the same causal certainty as a randomised experiment. It can still tell you which uncertainty is worth spending scarce traffic on.

Can you observe the result that actually matters?

Some outcomes arrive quickly. Someone clicks, registers or buys.

Others take longer.

A training business may care about whether it attracts people who are genuinely suitable and go on to complete the programme. A service business may care much more about profitable customers than cheap leads.

You can test an earlier proxy, but then you are relying on the proxy being meaningfully connected to the later result.

More applications are not automatically better if the extra applicants are a poor fit. A change can improve an early conversion metric while making the business worse off further down the journey.

A rigorous experiment on the wrong outcome can still produce a bad commercial decision.

Do you need a winner or an explanation?

Sometimes the question really is as simple as whether version A or version B performs better.

Imagine that version B changes the headline, page structure, proof and offer presentation at the same time. If it reliably outperforms A, the business may have enough information to choose B.

What it has not necessarily learned is why B won.

That matters if the purpose of the experiment is to understand something about customer behaviour that should guide other decisions. Changing several important things together makes that learning harder because the result belongs to the whole package.

If the business only needs to choose between two complete versions, that may be fine.

Can you measure both groups consistently?

High-consideration journeys often cross several systems and involve people as well as pages.

A prospect might move from a website into email, speak to someone in the team, pay through another system and receive support somewhere else. Important conversions may happen offline, and staff may make different decisions after leads arrive.

An experiment can look precise while the underlying journey is being observed inconsistently.

Before trusting the result, I want to know that both groups are being measured in a comparable way and that the outcome means the same thing on both sides.

If the measurement cannot support that comparison, fixing it may be more useful than launching another variant.

Sometimes the decision is already clear enough

If checkout is broken on mobile, fix it.

If an important eligibility condition on the website is factually wrong, correct it.

If customers cannot find their login because the link goes to the wrong place, repair the link.

Running an experiment would add very little to any of those decisions.

Testing is not free. It uses traffic, time and analytical attention, so additional evidence should have some realistic chance of changing what the business does.

Other evidence can make the eventual experiment better

Customer conversations, support messages, transaction records and behaviour elsewhere in the journey do not need to compete with experimentation.

They can make the experiment better.

“Conversion is weak” gives you almost nothing useful to test. Evidence from the journey can turn that into something more specific: people understand the offer but remain uncertain about one important condition, qualified prospects reach a particular step and disappear, or buyers repeatedly need a member of staff to explain the same issue.

At that point there is a clearer question to investigate.

If there is enough traffic, an A/B test may be an excellent way to answer it. If there is not, you can make a proportionate change and continue looking at the evidence without pretending you have experimental certainty.

Before I run a test, I want to know what uncertainty it is supposed to resolve and what I would do differently depending on the result. I also need enough relevant customers, an outcome I can trust and enough time to observe it.

If those conditions are not there, I would usually investigate the problem before building another variant.