How to diagnose a conversion problem before you optimise it

Steve New

I recently reviewed the sales journey for a substantial professional training programme. At first glance, there was plenty wrong with it.

Some prospects were still asking basic questions about participation and fit. There had been checkout and discount problems. Human follow-up occasionally failed. Onboarding generated support noise. A few people requested refunds shortly after the first live session. The sales material itself was long and had accumulated information over time.

You could make a reasonable case for changing almost any part of the journey.

That is exactly why the work needed diagnosis rather than a longer improvement list.

A good conversion audit can include serious quantitative and qualitative research, causal hypotheses and sensible prioritisation. The problem is not the word audit. The problem is treating the discovery of things that are weak, broken or improvable as sufficient evidence for what caused the commercial result and what should change.

Diagnosis is the reasoning that connects those observations to explanations, tests the alternatives and determines what action the evidence can support.

Conversion problems are diagnosis problems before they are optimisation problems.

In this case, the useful question was narrower than “What could we improve?”:

Where does a plausibly suitable prospect's progress first materially weaken between serious consideration and successfully starting the programme, why does it weaken, and what should change before the next sales cycle?

By plausibly suitable, I mean somebody whose goals, background and apparent ability to participate made them a credible fit independently of whether they eventually bought.

By the end of the analysis, the answer was much smaller than the original list of problems suggested.

The main repeated weakness was not initial interest, checkout or onboarding. It appeared when people who could already see what the programme offered had to decide whether joining was genuinely right and workable for them now.

Some of the difficulty at that point was avoidable. Some reflected real constraints that a good sales journey should preserve rather than try to overcome. Several obvious improvement projects turned out not to be justified at all.

None of the individual ingredients here is particularly exotic. Systematic CRO research already combines quantitative and qualitative evidence. Experience mapping starts by defining scope and intended use. Good customer research connects investigation to a business need. Causal-inference work forces alternative explanations into view. Decision analysis asks whether reducing an uncertainty is worth the cost of doing so. Controlled experimentation gives us stronger ways to estimate intervention effects when the conditions support it.1–6

What matters to me is how those ideas connect in an ambiguous commercial journey. The business decision determines which uncertainties matter. The evidence narrows the explanations that could change that decision. The consequence of being wrong determines how much confidence we need before acting.

What does the evidence support changing now?

Start with the decision you are trying to improve

“Conversion is weak” describes a result. It does not tell you what decision the business actually needs to make.

The real question might be whether to change the price, rebuild a sales page, improve acquisition, alter an offer, repair checkout, rethink onboarding, or simply stop spending time on an area that is already working well enough.

Those decisions require different evidence.

There are also two different decisions in the journey.

The business decision is what the organisation needs to decide or change. The customer decision is what the buyer is trying to work out.

In the training case, the business needed to know what to change before the next sales cycle. The customer was deciding whether a demanding six-month professional programme was sufficiently valuable, appropriate and practical for them to join now.

Once those are clear, the relevant population becomes clearer as well.

A visitor who spends thirty seconds on the page is not equivalent evidence to a prospect who attends an event, reads the offer, asks specific questions, considers the commitment and then decides not to join. Depending on the question, both may matter. They should not automatically be treated as the same kind of loss.

This is also why denominator problems matter.

Onboarding initially looked messy. There were support questions, access problems and inconsistent records, which made it easy to assume that buyers were getting lost after purchase.

Before interpreting those anecdotes, I reconstructed who was actually eligible to attend the first live Opening. The records contained duplicates, corrected identities, late purchasers and other cases that made the neatest initial population unreliable.

After reconciliation, 53 of 54 eligible participants attended at least part of the Opening.

That does not prove onboarding was excellent. It tells us nothing directly about prospects who never bought, and little about later persistence or success.

It does substantially weaken a narrower explanation: that widespread post-purchase onboarding failure was preventing buyers from reaching the first live session.

Without the right population, even a precise percentage can support the wrong diagnosis.

The broader principle is that the question determines which evidence matters. The fact that data exists does not.

Keep the observation separate from the explanation

Suppose twenty people start checkout and do not finish.

“Twenty people abandoned checkout” is an observation. “Checkout friction is causing lost sales” is an explanation.

Suppose several prospects mention price.

“Several prospects said the programme was expensive” is an observation. “The programme is overpriced” is an explanation.

The explanations may eventually prove useful. The mistake is behaving as though they arrived directly from the data.

The same observation can fit several explanations that would lead to very different actions. This is the same reason I treat disagreement between customer words and behaviour as diagnostic evidence rather than a vote between two sources.

Checkout abandonment might reflect a technical problem, unexpected price, loss of trust, a payment constraint, an unresolved fit question or somebody simply changing their mind.

A price objection might reflect genuine inability to pay, cash-flow timing, low priority, uncertainty about value, a poor comparison with alternatives or an entirely sensible decision that the programme is not worth the trade-off for that person.

I do not try to preserve every explanation that could conceivably be true. That becomes analysis for its own sake.

An explanation deserves to stay alive when it is plausible and when believing it rather than another live explanation would materially change what we do.

If several subtle explanations would all lead to the same cheap clarification on the same page, separating them may have little decision value.

If one explanation means lowering the price, another means repairing checkout, another means narrowing acquisition and another means leaving the journey alone because the loss is healthy qualification, the distinction matters.

Several explanations can also remain true at once. Diagnosis does not require one grand winner. It requires enough separation to know which mechanisms deserve which responses.

Look for evidence that separates the explanations

The training case began with several explanations that could all have sounded convincing in a meeting.

Perhaps prospects did not understand the programme's value. Perhaps the price was too high. Checkout had produced real problems. Onboarding looked noisy. The refund window raised the possibility that people were discovering a mismatch only after the programme began.

The useful evidence was not whatever happened to support one of those stories. It was evidence capable of changing their relative credibility.

Value was not simply a comprehension problem

Many serious prospects could describe what they valued in the programme: live practice, direct feedback, professional application, certification and the chance to develop practical capability.

That weakened a simple explanation that people were failing to progress because they did not understand what the programme offered.

It did not show that the perceived value was sufficient relative to the price, time or alternatives. Someone can understand an offer perfectly and still decide that the trade-off is not worthwhile.

The evidence narrowed the claim rather than eliminating the wider value question. A sweeping value-proposition rewrite became less attractive as the first project.

Price was real, but not one thing

There were genuine affordability cases. At least one plausibly suitable prospect could join once a workable payment arrangement was found, so financial constraints were clearly real in some purchases.

There was no clean population-level evidence showing that standard price explained most serious non-purchase.

“Price matters to some people” and “the standard price is the main conversion constraint” are different claims with very different consequences.

The first may justify payment options, scholarships or another route for particular customers. The second could justify changing the economics and positioning of every sale.

The available evidence supported the first much more strongly than the second.

Checkout contained defects without becoming the whole diagnosis

There had been genuine checkout and discount problems.

Those should be repaired.

What was missing was evidence that the standard checkout systematically prevented otherwise-ready buyers from purchasing.

Fix the demonstrated defect. Don't inflate its causal scope.

A broken payment path does not need a lengthy research project before repair. Its existence also does not establish that checkout explains the wider conversion problem.

Onboarding looked worse before the denominator was fixed

Onboarding produced some of the noisiest evidence in the journey.

Customers asked questions. Access and calendar issues occurred. Internal records were inconsistent enough that the apparent population itself needed reconstruction.

If I had stopped at anecdotal evidence, onboarding could easily have become a large improvement project.

Then the 53-of-54 Opening result changed the picture.

It did not erase the support problems. It changed what they meant.

Onboarding still contained things worth repairing, but the stronger explanation that buyers were failing in large numbers to reach the first live experience became difficult to defend.

This is the kind of evidence I find most useful in diagnostic work. It does not merely add another fact. It changes which actions still make sense.

Find the earliest defensible point where progress weakens

Once the main explanations have been narrowed, I want to know the earliest point the available evidence supports where a relevant customer's ability or willingness to make useful progress materially deteriorates.

I call that the first meaningful weakening.

It is not automatically the first measurable funnel drop and it is not necessarily the biggest percentage loss. It also does not automatically tell us what to fix first.

A problem can begin upstream while a separate downstream defect still deserves immediate attention. A journey can contain more than one meaningful constraint. Intervention priority still depends on consequence, prevalence, confidence, changeability and the downside of getting the action wrong.

The first meaningful weakening is primarily a way of locating the diagnostic chain, not a command to optimise the earliest possible point.

It also helps to distinguish three locations that are often collapsed into one. A contributing mechanism may begin in one place, customer progress may materially weaken later, and the visible symptom may appear later again. That is why a visible conversion problem can look bigger, later or more distributed than the underlying constraint.

A prospect might abandon at checkout while the important uncertainty began when an earlier page created a misleading expectation about price, workload or fit. Checkout is where the failure becomes visible. It does not follow that checkout created it.

In the training journey, the evidence pointed to a relatively narrow decision stage. Plausibly suitable prospects often understood what the programme offered, but still had to work out whether it was right and workable for them now.

Questions repeatedly involved experience level, workload, live attendance, certification, flexibility, professional fit, payment and who could give a reliable answer.

The earliest defensible weakening was therefore the move from:

“This looks valuable”

to:

“This is right and workable for me now.”

I would not claim that every prospect reached one precise timestamp where progress deteriorated. The evidence supported a stage of the decision, which was enough.

There is another question before calling that weakening a problem: should this customer actually progress?

Some people faced real constraints around workload, timing, family commitments or money. Better marketing could not and should not remove all of them.

Some friction also performed a useful qualification function. Participation requirements made the commitment more visible before purchase. A minimum financial contribution could act partly as a screen rather than simply an obstacle to maximise away. In one case, a prospect remained willing to buy while the pre-purchase interaction itself gave the business enough new information to decide that the longer relationship probably should not begin.

Calling something “healthy qualification” needs evidence. It cannot simply become a flattering explanation for poor conversion.

The loss should make sense in relation to real requirements, downstream fit, customer success or the kind of relationship the business is trying to create. This is why I treat non-purchase as something to classify before trying to eliminate it.

With that safeguard, I find four possibilities useful: avoidable loss, healthy qualification, an external constraint and genuinely unknown.

For a consequential purchase, the useful goal is not always the maximum number of immediate yeses. A good journey should help a suitable customer reach a sound decision, even when the answer is later or no.

Conversion diagnosis narrows observations and competing explanations through discriminating evidence to supported explanations, the first meaningful weakening, a decision threshold and proportionate action.
The diagnostic field narrows only as evidence removes explanations that no longer deserve action. The five outcomes are equally legitimate when supported by the evidence.

Decide whether the diagnosis is good enough to act on

A diagnostic can almost always be made more detailed. The useful question is whether more detail is worth acquiring.

By the end of the training analysis, I had fairly strong confidence about the location of the main repeated weakness. I had less confidence about exactly how much of the remaining difficulty came from price, workload, timing, schedule or fit.

More research could have narrowed those proportions. It was unlikely to change the first actions.

The recommendations were already relatively stable: make fit and feasibility easier to judge, maintain one reliable version of the current commercial and participation rules, and make serious buying questions much harder to lose in human follow-up.

Before stopping, I still want to challenge the emerging diagnosis.

What is the strongest remaining alternative? What evidence would make me materially change my mind? Is there contrary evidence I have explained away too easily?

Only after that adversarial check does the stopping rule become useful rather than an excuse to stop when the preferred answer feels comfortable.

My practical rule is to stop when the next action is stable enough that further investigation is unlikely to change it enough to justify the cost, delay or attention required. That is a practical application of the same idea formalised in decision analysis as the value of information: the value of learning depends on both the chance that current uncertainty could lead us to the wrong decision and the cost of being wrong.4

Two questions make the judgement easier: What decision could the next piece of evidence actually change? And how bad would it be if we acted now and were wrong?

Suppose the evidence moderately supports rewriting one paragraph that repeatedly leaves a serious prospect unsure about participation. The change is cheap, easy to reverse, affects a small part of the journey and will produce fresh feedback quickly.

I do not need extremely high confidence before trying it.

Now imagine the proposed response is to cut the standard price, rebuild the whole website, change the company's positioning or redesign the offer.

The same diagnostic confidence is not enough.

Those interventions affect more customers, cost more, consume more attention, can damage things that currently work and may take months before we learn whether they helped.

The evidence burden should therefore rise with the consequence of being wrong.

The intervention should not be bigger than the diagnosis.

That does not mean smaller interventions are intrinsically better.

Sometimes the underlying system genuinely needs major change. A platform may be unsupported, a product architecture may have become structurally incoherent, or repeated evidence may point to a problem that cannot economically be repaired with small adjustments.

In those cases a large intervention can be entirely proportionate.

The point is that the size of the action should follow the strength and scope of the diagnosis rather than the length of the improvement wish list. A website rebuild, for example, is a large bundled hypothesis unless the diagnosis actually supports that scope.

Decide what follows

Once a finding is strong enough to matter, there are several legitimate dispositions. A single diagnostic can produce several at once.

FIX

The problem is sufficiently established and the proportionate repair is clear.

If a payment button reliably fails, the failure can be reproduced, suitable buyers cannot proceed and the repair is straightforward, fix it.

The same applies to known contradictory information, a stale rule or another established problem whose repair does not depend on resolving a larger causal question first.

TEST

An important uncertainty remains, and deliberately changing something is a worthwhile way to learn about its effect.

Testing makes most sense when the alternatives are real, enough relevant exposure exists, the result can be observed in time to matter and different outcomes would actually change the decision. Controlled experiments can provide much stronger causal evidence than ordinary before-and-after observation, but only if the experiment and its metrics answer the commercial question that matters.6 In low-volume journeys, that is also why A/B testing can be the wrong tool even when it is technically possible.

The fact that something can be tested is not enough. A precise answer to a low-value question can still be a poor use of scarce customer attention.

MEASURE

A missing observation could materially change the action, and collecting it is worth doing before changing the underlying system.

In the next training cycle, useful measurement included recording why serious prospects defer or decline, separating genuine checkout failures from ordinary abandonment, tracking whether important questions receive a closed-loop answer, and capturing simple reasons for early refunds.

The point is not instrumentation for its own sake. It is obtaining a particular observation because that observation can resolve a live decision.

NOT YET

The issue may matter, but the contemplated action is not responsibly supported yet.

Several live explanations may still imply different large interventions. The relevant population may be impossible to reconstruct. The next natural cohort may produce much better evidence without forcing an artificial research exercise now.

NOT YET means the decision remains open.

LEAVE ALONE

Current evidence positively does not justify changing this.

The apparent problem may be healthy qualification. The area may already work adequately. The constraint may be external. The downside of changing it may exceed the likely benefit. Or another constraint may simply deserve the attention first.

That is different from uncertainty.

MEASURE means learn this now. NOT YET means do not make this decision yet. LEAVE ALONE means do not change this on the current evidence.

The training diagnostic produced several dispositions at once.

Specific defects deserved repair. Some unanswered questions needed better measurement. The evidence did not justify cutting the standard price, rebuilding checkout, overhauling onboarding, redesigning the curriculum, substantially increasing sales email or trying to eliminate every early refund.

Those non-actions were part of the value of the work.

For the larger recommendations, the diagnosis also had to remain revisable. I would change my view if better evidence showed that most plausibly suitable prospects actually failed to understand the offer, price overwhelmingly accounted for serious non-purchase, standard checkout failures were widespread, the first live experience repeatedly created the mismatch, or an earlier constraint consistently appeared before the fit-and-feasibility decision.

A useful diagnosis should be firm enough to act on and specific enough to be wrong.

Follow the intervention far enough to see whether it helped

The three main recommendations in this case were intended to make fit and feasibility easier to judge before purchase and reduce avoidable uncertainty when serious questions arose.

Their success should not be judged simply by whether the immediate purchase rate rises.

I would also want to know whether serious prospects reach decisions with fewer repetitive unresolved questions, whether fewer important conditions are discovered only after payment, whether suitable buyers still reach the first live experience successfully, and whether the pattern of early fit-correction refunds improves without suppressing useful qualification.

That is the more general principle.

A local conversion improvement matters only insofar as it helps the customer and business outcome the diagnosis was supposed to improve. Experimentation practice makes the same point through the choice of evaluation criteria: short-term metrics are useful when they provide a credible signal about the longer-term objective rather than merely being easy to move.6 This is also why I treat checkout as a useful event rather than the automatic commercial endpoint.

There is no need to track every possible downstream metric by default. Follow the result far enough that a local improvement cannot easily hide a worse outcome elsewhere, then update the diagnosis from what happened.

The professional training analysis never produced a complete explanation of every prospect's behaviour.

I could not say exactly what proportion of serious non-purchase came from price rather than workload, timing or fit. There was no single root cause that explained every support conversation, abandoned checkout and refund.

It did something more useful.

It made several expensive explanations less credible, identified the customer decision where progress repeatedly weakened, separated avoidable uncertainty from legitimate constraints, supported a small set of proportionate changes and gave us reasons not to start several much larger projects.

The business decision determined which uncertainty mattered. The evidence narrowed the explanations. The consequence of being wrong determined how far the investigation needed to go before action.

A good diagnosis reduces uncertainty until the next useful decision becomes clear.

Selected sources

  1. Peep Laja / CXL, Conversion Research and the ResearchXL approach.
  2. Jim Kalbach, Mapping Experiences, on framing, scope and experience mapping.
  3. Steve Portigal, Interviewing Users, 2nd edition, particularly the principle that research begins with a business need.
  4. Douglas W. Hubbard, How to Measure Anything, especially the treatment of the value of information.
  5. Nick Huntington-Klein, The Effect: An Introduction to Research Design and Causality, on research design, causal questions and alternative explanations.
  6. Ron Kohavi, Diane Tang and Ya Xu, Trustworthy Online Controlled Experiments, particularly experiment metrics and the Overall Evaluation Criterion.