CRO
A/B Testing
How to choose between A/B, split and multivariate testing, run the experiment honestly, and read the result without fooling yourself.
A/B, split and multivariate testing compared: which online experiment to run, how much traffic each needs, and how to avoid the mistakes that produce confident wrong answers.
What it is
An online glossaryExperimentAn experiment is a structured test used to evaluate hypotheses and measure outcomes.Open glossary term shows different glossaryVersionA version is a specific iteration of software or a product at a point in time.Open glossary term of the same experience to different users, then measures which performs better against a metric you agreed in advance.
The value is not the winning glossaryVariantA variant is a version of a design or experience used in testing or experimentation.Open glossary term. It is that a decision which would otherwise be settled by whoever is most senior in the room gets settled by what people actually did.
An experiment is only worth running if you would genuinely act on either result. If one outcome changes nothing, you are gathering reassurance rather than evidence.
A/B, split or multivariate: which experiment you need
The three are often used as though they were interchangeable. They are not, and the difference is mostly about how much you are changing and how much glossaryTrafficTraffic refers to the number of users visiting a website, app, or digital product over a given period.Open glossary term that costs you.
Key takeaway
A/B tells you whether one change helped. Split tells you which approach wins. Multivariate tells you how elements interact, and charges you a great deal of traffic for the privilege.
When to use it
glossaryExperimentAn experiment is a structured test used to evaluate hypotheses and measure outcomes.Open glossary term suit decisions that are reversible, measurable and worth the wait.
They are most useful when:
Testing is not a substitute for understanding. An experiment can tell you which of two options wins. It cannot tell you what the third, better option would have been.
How to run it
Set up properly
Be clear on the glossaryHypothesisA hypothesis is a testable assumption about how a change will impact an outcome.Open glossary term, the variations, and the single primary metric before anything is built. A test with several success metrics will always find one that moved.
Decide the glossarySample SizeSample size refers to the number of participants included in a study or test.Open glossary term and duration in advance, and commit to them. Stopping when the result looks good is the most common way A/B tests produce false conclusions.
Run the method
A/B testing compares two glossaryVersionA version is a specific iteration of software or a product at a point in time.Open glossary term under live conditions. Its authority comes entirely from the discipline around it. The same test run loosely produces confident nonsense.
- Split glossaryTrafficTraffic refers to the number of users visiting a website, app, or digital product over a given period.Open glossary term randomly and concurrently. Running one glossaryVersionA version is a specific iteration of software or a product at a point in time.Open glossary term this week and the other next introduces every difference between the two weeks.
- Change one thing, or accept that you will not know which change caused the result.
- Run for the planned period, covering whole weeks so weekday and weekend glossaryBehaviourBehaviour refers to how users interact with a system, including actions, patterns, and responses.Open glossary term are both represented.
- Collect the glossaryDataData is raw, uninterpreted information collected and stored so it can be analysed, processed, or used to inform decisions. On its own it carries no meaning; context and interpretation are what make it useful.Open glossary term without looking for a winner mid-flight. Repeated checking inflates the chance of a false positive substantially.
- Keep everything else constant (pricing, campaigns, glossaryReleaseA release is the point at which a product or feature is made available to users. It marks the transition from development to real-world use and often involves deployment, communication, and monitoring.Open glossary term) or the test measures those instead.
Focus on the discipline rather than the tooling. Most failed A/B programmes fail on stopping rules and glossarySample SizeSample size refers to the number of participants included in a study or test.Open glossary term, not on the glossaryPlatformA platform is a system or environment that enables users, services, or applications to interact, build, or operate.Open glossary term.
Capture and make sense of it
The value comes from a defensible answer. After the test, document:
- glossaryPerformancePerformance refers to how quickly and efficiently a system responds to user actions and processes tasks.Open glossary term against the pre-declared primary metric
- Whether the result is significant, and what the glossaryConfidence IntervalA confidence interval is a statistical range that estimates where the true value of a result is likely to fall.Open glossary term actually spans
- The winning glossaryVersionA version is a specific iteration of software or a product at a point in time.Open glossary term, including when the honest answer is no difference
- What the result implies for the next glossaryHypothesisA hypothesis is a testable assumption about how a change will impact an outcome.Open glossary term
Apply the learning, and record inconclusive tests too. A programme that only keeps its wins slowly convinces itself of things that are not true.
Where it goes wrong
The failures are consistent and mostly self-inflicted.
What you get from it
Done properly, this method gives you:
Key takeaway
It replaces the loudest opinion with a number.
Get in touch
If this sounds like something you need, we can help you design experiments that answer the question you actually have.
No guesswork. No assumptions. Just clarity you can build on.
FAQ
Common questions
A few practical answers to the questions that usually come up around this method.
What is A/B testing?
It is an experiment where two versions of the same experience are shown to different users, and performance is measured against a metric agreed in advance.
What is the difference between A/B testing and split testing?
A/B testing usually varies one element on the same page. Split testing compares substantially different designs, often hosted at separate URLs. The terms are frequently used interchangeably, but the size of the change is what actually differs.
When should you use multivariate testing instead?
When you need to know how several elements interact rather than whether one change helped, and when you have enough traffic to give every combination its own sample. On most sites that condition is not met.
How much traffic do you need?
Enough to detect the effect you would care about, which depends on your baseline conversion rate. Calculate the sample size before you launch rather than running until the answer looks right.
Does A/B testing improve UX?
It improves decisions, which is not the same thing. It tells you which of the options you thought of performs better. Research is what tells you whether you thought of the right options.
Quick take
If you want to know what actually works rather than what everyone thinks works, run an experiment. Which kind depends on how big the change is and how much traffic you have.
Related Services



