Key Takeaways
- Fixing and optimising are two different jobs: fixing is remediating what is clearly broken or wrong (do it now, on judgement); optimising is a rigorous experiment program that improves a working page over time.
- Do not A/B test the obvious — repair broken forms, mismatched messages, bloated forms, and slow pages immediately; save experiments for genuinely uncertain improvements.
- Optimise from research, not opinion: gather the inputs (analytics, session replay, customer feedback) so your tests target real problems, then prioritise with a framework like ICE or PIE.
- A valid A/B test changes one variable, runs to a pre-set sample size and duration, and reaches genuine statistical significance — most 'wins' are false positives from peeking and stopping early.
- Interpret honestly: check segments, watch for results that reverse under segmentation, and be willing to conclude a test was flat rather than shipping noise as a win.
- If your page lacks the traffic to run valid tests (most do), improve it with research-backed best-practice changes rather than underpowered experiments that manufacture false confidence.
Fixing and Optimising Are Two Different Jobs
The phrase 'fix and optimise' contains two jobs that most teams blur together, and separating them is the first and most useful move. Fixing is remediation: repairing something that is broken or clearly wrong — a form that fails on mobile, a page whose message does not match its traffic, a form asking for ten fields when it needs three, a page that loads slowly. Optimising is different: it is the disciplined process of improving an already-working page through experiments, testing genuinely uncertain changes to squeeze out gains over time. These jobs have different methods, different speeds, and different tools, and using the wrong one for the situation is a core reason CRO underperforms — teams run months of A/B tests on a page that has an obvious broken form, or they ship untested guesses to a page that needed rigorous experimentation.
The rule that separates them is simple: you fix what is clearly wrong immediately, and you optimise what is genuinely uncertain with experiments. You do not A/B test whether a broken form should work, whether a page should load fast, whether a message should match its ad, or whether a form should stop asking for unnecessary fields — these are known-good practices and known problems, and testing them wastes time and traffic while the loss continues. You do test whether one clear headline beats another, whether a different offer framing lifts conversions, whether a shorter page outperforms a longer one — changes where reasonable people disagree and the answer is not obvious. Confusing the two produces two failures: treating obvious fixes as tests (slow, wasteful) and treating uncertain changes as obvious fixes (shipping guesses that may hurt). Know which job you are doing.
This guide covers both, in order, because the sequence matters. You fix the clearly-wrong things first — often the issues surfaced by diagnosing your bottleneck levers or running a root cause analysis on a drop — because there is no point running delicate experiments on a page with a broken form, and remediating the obvious problems both stops the bleeding and gives you a clean baseline to optimise from. Then you run a real optimisation program on the improvements that are genuinely uncertain. Most of this guide is about that second job, because remediation is mostly a matter of judgement and doing the obvious, while optimisation is where rigour is required and where most teams go wrong — running experiments that look scientific but are statistically invalid, and shipping false winners with confidence.
Fix First: Remediate the Obvious Before You Optimise
Before any experimentation, remediate the clearly-wrong things, because they are certain losses you can stop immediately without the time and traffic cost of a test. The remediation list comes straight from diagnosing your page — the bottleneck levers that increase conversion rate, and the root cause if a drop triggered this work. If your message does not match your traffic, fix it now. If your form asks for unnecessary fields, remove them now. If your page is slow on mobile, fix the performance now. If a form or script is broken, repair it now. None of these require a test to justify, because they are known problems with known-good fixes, and the only thing a test would add is a delay during which the loss continues. Remediation is fast, high-return, and the right first move.
The reason to fix before you optimise, beyond stopping obvious losses, is that a broken or clearly-wrong page corrupts the experiments you would run on it. If you A/B test a new headline on a page whose form is broken on mobile, the technical break dominates the result and drowns out the effect you are trying to measure — you learn nothing about the headline because the page's real problem is elsewhere. Remediation gives you a clean, working baseline where experiments can actually detect the effects of the changes you test, rather than being swamped by uncorrected problems. Optimising a broken page is like tuning an engine with a hole in the fuel line; fix the obvious defects first, then tune.
There is a judgement call at the boundary between fixing and optimising, and it is worth naming: some changes are obvious enough to just make, and some are uncertain enough to test, and the line depends on your confidence and your traffic. A change you are highly confident improves things, that follows established best practice, and that carries little risk, you can often just make — especially if your traffic is too low to test it validly anyway. A change where reasonable people disagree, that could plausibly hurt, or that is expensive to build, deserves a test if you have the traffic to run one. The remediation phase is for the high-confidence, known-problem fixes; the optimisation phase is for the genuinely uncertain improvements. When in doubt, ask whether you would be comfortable shipping the change without a test — if yes, it is a fix; if no, it is a test.
The Optimisation Loop: Research, Prioritise, Test, Learn
Optimisation proper is a loop, not a series of one-off tests, and running it as a disciplined cycle is what separates a CRO program that compounds from a scatter of random experiments. The loop has clear stages: research to understand where and why you lose conversions, prioritise the ideas so you work on what matters most, form a genuine hypothesis, design and run a valid experiment, analyse the result honestly, and implement or iterate — then repeat, each cycle informed by the last. The exploded view below lays out the loop stage by stage, with what each stage requires to be done well, because the value of the program is in doing every stage properly, and most teams skip straight from a vague idea to a poorly-run test, missing the research, prioritisation, and rigour that make the loop actually work.
The continuous loop of a real conversion optimisation program. One, research: gather the inputs — analytics for where, session replay for why, voice-of-customer for the objection — so tests target real problems, not opinions. Two, prioritise the resulting ideas with a framework (ICE, PIE, or evidence-tied PXL) and work the queue top-down, scoring honestly. Three, form a genuine hypothesis naming the observed problem, the change, and the expected effect with a reason, so every outcome is informative. Four, design and run a valid experiment — one variable, a pre-calculated sample size, one to two full weekly cycles, genuine statistical significance, and no peeking or stopping early, which manufacture false winners from noise. Five, interpret honestly: check the result by segment, watch for reversals, and have the courage to call a test flat rather than cherry-pick a favourable segment. Six, implement winners and feed losses and flats back as sharper hypotheses, building an accumulating tested model of what moves your visitors. And the reality most advice ignores: if your traffic is too low to reach significance — true for most pages — improve on research-backed judgement and best practice rather than underpowered tests that lie to you.
The first stage, research, is the one most often skipped and the one that determines whether everything downstream is worthwhile. Optimising from opinion — 'I think we should try a green button', 'let's test a shorter page' — produces a stream of tests untethered from real problems, most of which are flat because they were solving problems the visitors do not have. Optimising from research means gathering the inputs first: analytics to see where conversions leak, session replay and heatmaps to see why, and voice-of-customer feedback to hear the objection, so your tests target real, observed problems. A single test grounded in a genuine, observed problem is worth ten tests based on internal guesses, because it is aimed at something that is actually costing you conversions. Research is what makes optimisation efficient; without it, you are running experiments at random and hoping.
Research produces more ideas than you can test, which is why the second stage, prioritisation, matters — you cannot test everything, so you must work on the highest-value ideas first. This is where a scoring framework earns its place, turning a pile of ideas into a ranked queue rather than a debate about whose idea to try. The next section covers the frameworks; the point here is that a real optimisation program is a disciplined loop — research, prioritise, hypothesise, test, learn — run continuously, with each stage done properly. Skip the research and you test the wrong things; skip the prioritisation and you test low-value things; skip the rigour in testing and you ship false winners. The loop only works when every stage does.
Prioritise: Decide What to Work On With a Framework
Research surfaces more opportunities than you have the traffic or time to test, so the discipline of prioritisation is what stops a program from either testing trivial ideas or arguing endlessly about what to try. A scoring framework makes prioritisation explicit and honest: you score each idea on a few dimensions, rank by the total, and work the queue top-down, which both surfaces the highest-value ideas and removes the politics of whose idea gets tested. The most common frameworks are ICE and PIE, and a more evidence-oriented one called PXL, and the table below compares what each scores so you can pick one that fits how your team thinks.
| Framework | What it scores | Best for | Watch out for |
|---|---|---|---|
| ICE | Impact, Confidence, Ease (each 1–10) | Fast, lightweight prioritisation | Scores are subjective; confidence is easily inflated |
| PIE | Potential, Importance, Ease | Weighing a page's potential and traffic importance | Similar subjectivity; 'importance' overlaps with impact |
| PXL | Structured yes/no questions tied to evidence | Reducing bias by grounding scores in data | More setup; heavier to run |
The choice of framework matters less than using one consistently, because the value is in forcing an explicit, comparable score for every idea rather than deciding by loudest voice or most senior opinion. ICE is the most popular for its simplicity — score impact, confidence, and ease, and rank by the product or sum — and it is a fine default. PIE is similar with a slightly different emphasis on a page's potential and its traffic importance. PXL, favoured by more rigorous teams, replaces subjective 1–10 scores with structured questions tied to evidence (is this change above the fold, is it addressing an issue found in research, how big is the change), which reduces the tendency to inflate scores for ideas you personally like. Whichever you use, the honest trap to avoid is gaming your own confidence scores to justify testing a pet idea — the framework only helps if you score honestly.
Prioritisation also enforces a healthy bias toward high-impact, backed-by-research changes over low-impact tweaks, which is the antidote to the button-colour-test trap. A framework that weights impact and evidence naturally ranks a test that addresses a major, research-confirmed problem above a cosmetic tweak that would move nothing even if it 'won'. This is how prioritisation quietly improves your whole program: it does not just decide the order, it steers effort toward the changes that can actually move the number, and away from the trivial tests that fill CRO backlogs and produce flat results. Score honestly, work the queue top-down, and you spend your limited testing capacity where it counts.
Form a Real Hypothesis, Not a Random Tweak
Before designing a test, you need a genuine hypothesis, because a hypothesis is what makes a test a learning exercise rather than a coin flip. A real hypothesis has three parts: the problem you observed, the change you are making, and the effect you expect, with a reason. 'Because session recordings show visitors abandoning at the phone-number field (problem), removing the phone-number field (change) will increase form completion (expected effect), because it removes friction from an action people already want to take (reason).' Contrast that with 'let's try removing the phone field and see what happens', which is a tweak, not a hypothesis — it has no observed problem behind it and no reasoning, so whatever the result, you learn little you can generalise.
The value of a real hypothesis is that it makes every outcome informative. If a well-formed hypothesis wins, you have confirmed a specific belief about your visitors that you can apply elsewhere — friction at required fields costs completions, so look for it across your forms. If it loses, you have learned that your model of the visitor was wrong in a specific way, which is also valuable. A random tweak that wins tells you only that this particular tweak happened to help this particular page, with no reasoning to transfer; a random tweak that loses tells you nothing at all. Because your testing capacity is limited — most pages can run only a handful of valid tests at a time — every test should be a hypothesis that teaches you something regardless of outcome, not a lottery ticket. The hypothesis is what converts a test from a gamble into a unit of knowledge.
Forming good hypotheses also disciplines the research-to-test pipeline, because a hypothesis must name an observed problem, which forces you to have actually done the research. If you cannot state the problem your test addresses — with evidence from analytics, replay, or feedback — then you do not have a hypothesis, you have a hunch, and it belongs back in the research phase, not in the test queue. This is a useful filter: requiring a real hypothesis for every test automatically screens out the opinion-based tweaks that clog CRO programs and produce flat results. When your backlog is full of proper hypotheses, each tied to a real problem and a clear expected effect, your program is aimed at things that matter and set up to learn from every outcome.
Design and Run a Valid Experiment — Where Most CRO Goes Wrong
This is the stage where most CRO programs quietly fail, because running a statistically valid experiment is harder than the tools make it look, and an invalid test produces confident, precise, wrong conclusions that are worse than no test at all. The tool splits your traffic and reports a winner; it does not guarantee the test was valid, and the validity is your responsibility. Four things determine whether an A/B test tells you the truth, and the table below lays them out as a checklist to run before you trust any result. Get these wrong and your 'winner' is very likely a false positive — a random fluctuation the tool crowned because you stopped at the right moment.
| Validity requirement | What it means | The failure it prevents |
|---|---|---|
| One variable | Change only one thing between control and variant | Not knowing which change caused the result |
| Adequate sample size | Enough visitors and conversions per variant, set before you start | Concluding from noise on too few conversions |
| Full duration | Run at least one to two full business cycles (weeks, not days) | Being fooled by a good day or a weekday/weekend skew |
| Statistical significance | Reach a genuine significance threshold set in advance | Shipping a random fluctuation as a win |
| No peeking / early stopping | Do not check repeatedly and stop when it looks good | Manufacturing false winners by stopping at a lucky moment |
The two failures that cause the most damage are peeking and stopping early, and they are so common they deserve special warning. A conversion rate bounces around during a test, and if you watch it continuously and stop the moment the variant is ahead with an apparently significant result, you will 'win' tests that are actually flat — because with enough looks, random noise will cross the significance line at some point by chance, and you stopped exactly there. This is not a subtle statistical nicety; it is the single most common way real teams ship changes that do nothing, believing they proved they worked. The discipline is to set your sample size and duration in advance, run the test to completion regardless of what the number does in the middle, and only evaluate at the end. Tools that offer 'always-valid' or sequential statistics can mitigate this, but the safest habit is to decide the stopping rule before you start and hold to it.
Adequate sample size and full duration are the other two you cannot fudge. A test needs enough conversions per variant to distinguish a real effect from noise — not enough visitors, but enough conversions — and the smaller the effect you want to detect, the more you need, which is why detecting a tiny lift requires enormous traffic. And a test must run long enough to cover the natural cycles of your traffic — at least one to two full weeks — because conversion behaviour differs by day of week and by the mix of traffic across a cycle, and a test run for three days may be measuring a weekday skew rather than a real effect. Calculate the required sample size before you start (there are standard calculators for this), commit to running until you reach it across at least a full cycle, and do not evaluate before then. This rigour is unglamorous and it is exactly what separates CRO that produces real, durable gains from CRO that ships a stream of false winners and wonders why the aggregate conversion rate never moves.
Interpret Honestly: Segments, Reversals, and the Courage to Call It Flat
A completed, valid test still has to be interpreted honestly, and honesty here means a few specific disciplines that resist the pull to declare victory. The first is to check the result by segment, because an overall result can hide important differences — a variant that wins overall might lose badly on mobile while winning big on desktop, which changes what you should actually ship (perhaps the desktop version for desktop only). Segmentation can also reveal a result that reverses: a change that appears to win in aggregate but loses within every individual segment, an artefact of shifting segment mix during the test. Looking at the result by device, source, and new-versus-returning visitor before you ship protects you from shipping a change that helps in the summary but hurts the visitors you care about.
The second discipline is the courage to call a test flat, which is harder than it sounds because there is organisational pressure to show that tests produce wins. Many valid tests show no significant difference, and the honest, correct conclusion is that the change did not matter — not that you should squint at the data until you find a favourable segment to declare victory on, which is called cherry-picking and is how false positives get shipped. A flat test is a real result: it tells you that the change you believed would help did not, which refines your model of your visitors and saves you from shipping something that does nothing. A CRO program that only ever reports wins is not a rigorous program; it is one that is fooling itself, because real testing produces plenty of flat and losing results, and treating those honestly is what keeps the wins real.
The third discipline is to treat losses and flats as learning, not failure, and to feed them back into the loop. A test that loses has taught you that your hypothesis about the visitor was wrong in a specific way, which is genuinely useful information that improves your next hypothesis — that is why forming real hypotheses matters, because it makes even losses informative. Over many cycles, a disciplined program accumulates a model of what actually moves your visitors, and that compounding understanding is the real asset a CRO program builds, more than any single winning test. The teams that improve conversion rate durably are not the ones with the most 'wins' on a slide; they are the ones who run valid tests, interpret them honestly including the flats and losses, and let the accumulated learning sharpen every subsequent test.
The Reality Most Advice Ignores: What If You Can't Run Valid Tests?
Here is the truth most CRO advice omits: the majority of landing pages do not have enough traffic to run valid A/B tests in any reasonable timeframe, and pretending otherwise leads teams to run underpowered tests that produce false winners and worse decisions than making no changes at all. Reaching statistical significance requires a certain number of conversions per variant, and a page with modest traffic and a normal conversion rate may need months to accumulate that for even a single test — during which the world changes and the test is contaminated. Running a two-week test on such a page and declaring a winner is not science; it is noise dressed as evidence, and acting on it is worse than not testing, because it builds false confidence. Ignoring this reality is one of the biggest failures of popular CRO advice.
So what do you do when you cannot run valid tests, which is most of the time for most pages? You optimise by judgement grounded in research and best practice, rather than by invalid tests. This means: fix the clearly-wrong things (which never needed a test anyway), make research-backed changes you are confident in based on the evidence from analytics, session replay, and customer feedback, and follow established conversion best practices — clear message match, low friction, strong proof, one CTA, fast pages — because these are supported by broad evidence even if you cannot prove them on your specific low-traffic page. You accept that you are improving on informed judgement rather than page-specific proof, which is the honest and correct approach when the traffic for proof does not exist. This is not a lesser method; it is the appropriate method for the traffic you have, and it produces real improvement without the self-deception of underpowered testing.
If experimentation is genuinely important to your growth and your traffic is the constraint, there are two honest paths: concentrate your testing on the highest-traffic pages where valid tests are possible and apply the learnings elsewhere, or focus on larger, bolder changes (which need less traffic to detect than small tweaks) rather than the marginal optimisations that require enormous samples. What you should not do is run a stream of underpowered tests and act on their false winners, because that is optimisation theatre that makes worse decisions than honest judgement would. The whole point of this guide is to help you actually increase conversion rate, and that means fixing what is broken immediately, testing what is genuinely uncertain when you have the traffic to test it validly, and improving on research-backed judgement when you do not — rather than performing the rituals of CRO while fooling yourself about the results. If you want help building a real optimisation program sized to your actual traffic, or the owned measurement that makes both testing and judgement trustworthy, that is exactly the kind of work our team does with growth teams.
Frequently Asked Questions
- What is the difference between fixing and optimising conversion rate?
- Fixing is remediation — repairing what is broken or clearly wrong, like a form that fails on mobile, a message that does not match its traffic, a form with unnecessary fields, or a slow page. You do these immediately on judgement, because they are known problems with known-good fixes and testing them just delays the fix while the loss continues. Optimising is different: it is a disciplined program of experiments that improve an already-working page over time, testing genuinely uncertain changes where reasonable people disagree and the answer is not obvious. The rule that separates them is whether you would be comfortable shipping the change without a test — if yes (a broken form, an obvious best practice), it is a fix; if no (one headline versus another, a different offer framing), it is a test. Fix the clearly-wrong things first to stop losses and get a clean baseline, then optimise the uncertain improvements with real experiments.
- How do I prioritise which conversion tests to run?
- Use a scoring framework so prioritisation is explicit and honest rather than a debate about whose idea to try. ICE scores each idea on Impact, Confidence, and Ease; PIE scores Potential, Importance, and Ease; and PXL replaces subjective 1–10 scores with structured, evidence-tied yes/no questions to reduce bias. Score every idea, rank by the total, and work the queue top-down. The choice of framework matters less than using one consistently and scoring honestly — the common trap is inflating your own confidence scores to justify a pet idea. A good framework naturally steers effort toward high-impact, research-backed changes over cosmetic tweaks, which is the antidote to the button-colour-test trap. Prioritisation is essential because research surfaces more opportunities than you can test, so you must spend your limited testing capacity on the ideas that can actually move the number.
- Why do my A/B tests keep showing wins that don't hold up?
- Almost always because the tests are statistically invalid, and the most common culprits are peeking and stopping early. A conversion rate fluctuates during a test, and if you watch it continuously and stop the moment the variant is ahead with an apparently significant result, you will crown random noise as a winner — with enough looks, chance alone will cross the significance line, and you stopped exactly there. The fix is to set your sample size and duration in advance, run the test to completion regardless of what the number does mid-test, and only evaluate at the end. The other common failures are too small a sample (not enough conversions per variant to distinguish a real effect from noise), too short a duration (fewer than one to two full weeks, so you measure a weekday skew), and changing more than one variable at once. Run through a validity checklist before trusting any result, because a false winner is worse than no test — it ships a change that does nothing with full confidence.
- What if my landing page doesn't have enough traffic to A/B test?
- This is the reality for most pages, and the honest answer is to optimise by research-backed judgement rather than by invalid tests. Reaching statistical significance requires a certain number of conversions per variant, and a modest-traffic page may need months to accumulate that for even one test — so running a two-week test and declaring a winner is noise dressed as evidence, and acting on it is worse than not testing. Instead: fix the clearly-wrong things (which never needed a test), make changes you are confident in based on your analytics, session replay, and customer feedback, and follow established best practices like message match, low friction, strong proof, one CTA, and fast pages, which are supported by broad evidence even without page-specific proof. If testing is genuinely important, concentrate it on your highest-traffic pages where valid tests are possible, or test larger, bolder changes that need less traffic to detect — but never run a stream of underpowered tests and act on their false winners.
- How long should I run an A/B test?
- Long enough to reach a pre-calculated sample size and to cover at least one to two full business cycles — in practice, usually a minimum of one to two weeks, and often longer for lower-traffic pages. The two constraints are sample size and duration, and both must be satisfied. Sample size: you need enough conversions per variant to distinguish a real effect from noise, which you calculate before starting based on your baseline conversion rate and the smallest lift you want to detect — smaller effects require far more conversions. Duration: even if you hit the sample size quickly, run the test across full weekly cycles, because conversion behaviour differs by day of week and by the traffic mix across a cycle, so a three-day test may measure a weekday skew rather than a real effect. Decide both the sample size and the minimum duration in advance, run to completion without stopping early when the number looks good, and only then evaluate the result.